Model garden Router · on request Dedicated · available

paza Phi 4 multimodal instruct

Runs as your own dedicated GPU deployment, ready to launch from the wizard. Data stays in Europe.

Fine-tuning was performed on the entire unified multilingual ASR dataset, comprising the mentioned six languages, to encourage cross-lingual generalization. During fine-tuning, only the audio-specific components: audio embedding module, audio encoder, and audio projection layers,...

microsoft/paza-Phi-4-multimodal-instruct On request
audio->text · microsoft · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
5.6B
Parameters
131K
Context window
32GB
Minimum VRAM
POST /api/v1/audio/transcriptions On request

Specifications

Parameters 5.6B
Context window 131,072 tokens
Minimum VRAM 32 GB
Architecture Phi4MMForCausalLM (vLLM)
License mit
Modality audio->text
Released January 2026
Publisher microsoft ↗

Pricing

Shared router · per token
On request
Not available on the shared router. Pricing on request as a dedicated GPU deployment.
Dedicated GPU · per hour
from €1,95 per hour
Your own vLLM instance on European cloud (32 GB VRAM), billed hourly.

Shared EU router, pay-per-token, scale-to-zero. Dedicated GPU deployments are billed hourly, see pricing.

Call it now

Drop-in replacement for OpenAI: change only the base URL and API key. The Anthropic format (/v1/messages) is supported too.

curl https://hostyourai.com/api/v1/audio/transcriptions \
  -H "Authorization: Bearer hyai-..." \
  -F model="microsoft/paza-Phi-4-multimodal-instruct" \
  -F file="@meeting.mp3"

Frequently asked questions

Can I run paza Phi 4 multimodal instruct in the EU?

Yes. HostYourAI runs paza Phi 4 multimodal instruct on GPUs in European datacenters via vLLM. Prompts and outputs never leave the EU and there is no US cloud provider in the chain.

Is hosting paza Phi 4 multimodal instruct GDPR-compliant?

Yes. All processing happens inside the EU, a Data Processing Agreement (DPA) is available and the subprocessor list is public. Open-source weights also mean: no training on your data.

How much does paza Phi 4 multimodal instruct cost?

paza Phi 4 multimodal instruct needs several GPUs at once, so it runs as a dedicated deployment billed per GPU-hour rather than per token. Tell us your volume and we will work it out with you.

Is the API OpenAI-compatible?

Yes. You use the standard OpenAI SDKs with a custom base URL (https://hostyourai.com/api/v1). The Anthropic Messages API is supported as a drop-in as well.

More models from Microsoft

Fara1.5 27B

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

27B 262K context View model →
Fara1.5 4B

Fara1.5-4B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

4.5B 262K context View model →
GELab Zero 4B preview Sico Evolution

GELab Zero 4B preview Sico Evolution is an multimodal language model from Microsoft with 4.4B parameters, hosted on EU GPUs via an OpenAI-compatible API.

4.4B View model →
Fara1.5 9B

Fara1.5-9B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

9.4B 262K context View model →
harrier oss v1 27b

harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.

27B 131K context View model →
harrier oss v1 270m

harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.

0.3B 33K context View model →

Try paza Phi 4 multimodal instruct for free

Creating an account takes a minute. Your API key works with paza Phi 4 multimodal instruct right away.

Start for free