microsoft

microsoft

Phi-4-multimodal-instruct

A lightweight 5.6B parameter model unifying text, vision, and speech in a single neural network for cross-modal reasoning, speech recognition, and image understanding.

Available on FriendliAI

Dedicated Endpoints

Run this model inference on single tenant GPU with unmatched speed and reliability at scale.

Learn more

Model Details

Model Provider

microsoft provider

microsoft

Model Tree

Base
this model

Input Modalities

TextAudioImage

Output Modalities

Text

Supported Functionality

Dedicated Endpoints

Explore FriendliAI today