Model garden Router · op aanvraag Dedicated · beschikbaar

Phi 4 mini flash reasoning

Dit model draait als eigen dedicated GPU-deployment, direct te starten via de wizard. Data blijft in Europa.

Phi-4-mini-flash-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 64K token context length.

microsoft/Phi-4-mini-flash-reasoning Op aanvraag
text->text · microsoft · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
3.9B
Parameters
262K
Contextvenster
16GB
Minimale VRAM
POST /api/v1/chat/completions Op aanvraag

Specificaties

Parameters 3.9B
Contextvenster 262,144 tokens
Minimale VRAM 16 GB
Architectuur Phi4FlashForCausalLM (vLLM)
Licentie mit
Modaliteit text->text
Uitgebracht June 2025
Uitgever microsoft ↗

Prijzen

Gedeelde router · per token
Op aanvraag
Niet beschikbaar op de gedeelde router. Prijs op aanvraag als dedicated GPU-deployment.
Dedicated GPU · per uur
vanaf €2,22 per uur
Eigen vLLM-instance op Europese cloud (16 GB VRAM), per uur afgerekend.

Gedeelde EU-router, pay-per-token, scale-to-zero. Dedicated GPU-deployments worden per uur afgerekend, zie prijzen.

Direct aanroepen

Drop-in vervanger voor OpenAI: wijzig alleen de base-URL en de API-key. Ook het Anthropic-formaat (/v1/messages) wordt ondersteund.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "microsoft/Phi-4-mini-flash-reasoning",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Veelgestelde vragen

Kan ik Phi 4 mini flash reasoning in de EU draaien?

Ja. HostYourAI draait Phi 4 mini flash reasoning op GPU's in Europese datacenters via vLLM. Prompts en outputs verlaten de EU niet en er is geen Amerikaanse cloudprovider in de keten.

Is Phi 4 mini flash reasoning hosten AVG/GDPR-compliant?

Ja. Alle verwerking vindt plaats binnen de EU, er is een verwerkersovereenkomst (DPA) beschikbaar en de subprocessor-lijst is openbaar. Open-source gewichten betekenen ook: geen training op jouw data.

Wat kost Phi 4 mini flash reasoning?

Phi 4 mini flash reasoning is beschikbaar op aanvraag: via de gedeelde EU-router na kort overleg, of als eigen dedicated GPU-deployment per uur. Vertel ons je use case, dan zetten we het voor je klaar.

Is de API compatibel met OpenAI?

Ja. Je gebruikt de standaard OpenAI-SDK's met een aangepaste base-URL (https://hostyourai.com/api/v1). Ook de Anthropic Messages API wordt ondersteund als drop-in.

Andere modellen van Microsoft

Fara1.5 27B

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

27B 262K context Bekijk model →
Fara1.5 4B

Fara1.5-4B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

4.5B 262K context Bekijk model →
GELab Zero 4B preview Sico Evolution

GELab Zero 4B preview Sico Evolution is een multimodaal taalmodel van Microsoft met 4.4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.

4.4B Bekijk model →
MagenticBrain

MagenticBrain is a 14B-parameter orchestration model from Microsoft Research AI Frontiers. It plans multi-step tasks, calls declared tools, and coordinates sub-agents. It does not execute actions itself — every real-world side effect happens inside a host harness.

15B 41K context Bekijk model →
Fara1.5 9B

Fara1.5-9B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

9.4B 262K context Bekijk model →
harrier oss v1 27b

harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.

27B 131K context Bekijk model →

Probeer Phi 4 mini flash reasoning gratis

Account aanmaken duurt een minuut. Test Phi 4 mini flash reasoning direct in de playground.

Start gratis