HostYourAI offers three execution modes side by side on a single account and credit balance. Start with EU Hosted inference, move to dedicated capacity when the workload or compliance profile requires it.
1. EU Hosted Gateway: pay per token
One OpenAI-compatible API key, model catalog, hyai/auto, and scale-to-zero shared capacity. Best for SaaS integrations, agencies, agent apps, and experimentation. De tarieven per model (EUR per miljoen tokens), exact de prijzen waarmee je verbruik wordt afgerekend:
Chat- en visionmodellen
| Model | Context | Invoer €/M | Cached €/M | Uitvoer €/M |
|---|---|---|---|---|
| Pixtral 12B | 125K | 0.23 | 0.23 | |
| Qwen3.5 9B | 256K | 0.17 | 0.05 | 0.23 |
| Qwen2.5 14B Instruct AWQ | 32K | 0.08 | 0.25 | |
| DeepSeek V4 Flash | 1M | 0.29 | 0.07 | 0.35 |
| Mistral Small 3.2 24B | 125K | 0.17 | 0.40 | |
| Gemma 3 27B | 40K | 0.29 | 0.58 | |
| DeepSeek V3.2 | 160K | 0.35 | 0.09 | 0.58 |
| GPT-OSS 120B | 128K | 0.16 | 0.63 | |
| Qwen3 235B A22B Instruct | 256K | 0.21 | 0.63 | |
| Qwen2.5 VL 72B | 125K | 0.26 | 0.79 | |
| Qwen3 Coder 30B A3B | 128K | 0.23 | 0.92 | |
| Llama 3.3 70B Instruct | 128K | 1.04 | 1.04 | |
| MiniMax M2.5 | 192K | 0.35 | 0.09 | 1.38 |
| MiniMax M3 | 200K | 0.46 | 0.12 | 2.30 |
| DeepSeek R1 0528 | 160K | 0.76 | 0.20 | 2.99 |
| Kimi K2.5 | 256K | 0.58 | 0.15 | 3.22 |
| GLM 5 | 198K | 1.15 | 0.29 | 3.68 |
| Qwen3.5 122B A10B | 256K | 0.58 | 0.15 | 4.03 |
| DeepSeek V4 Pro | 1M | 2.01 | 0.51 | 4.03 |
| Qwen3.5 397B A17B | 256K | 0.69 | 4.14 | |
| Kimi K2.6 | 256K | 1.15 | 0.29 | 4.60 |
| GLM 5.1 | 198K | 1.48 | 4.66 | |
| GLM 5.2 | 1M | 1.73 | 0.44 | 5.18 |
| Kimi K2.7 Code | 256K | 1.44 | 0.36 | 5.18 |
| Mistral Medium 3.5 | 128K | 1.73 | 8.63 | |
| Kimi K3 | 1M | 3.45 | 0.86 | 17.25 |
Embeddings
| Model | Context | Invoer €/M tokens |
|---|---|---|
| Qwen3 Embedding 8B | 32K | 0.12 |
| BGE Multilingual Gemma2 | 8K | 0.12 |
Embeddings hebben geen output-tokens; je betaalt alleen voor de input.
Audio naar tekst
| Model | Prijs per minuut audio |
|---|---|
| Whisper Large v3 | € 0,0035 |
Afgerekend per seconde audio, niet per minuut afgerond.
Spraak-, beeld- en videomodellen zijn op aanvraag beschikbaar via /contact.
Modellen zonder gedeelde tokenprijs draaien als dedicated deployment en worden per project geoffreerd via /contact.
EU Hosted means the request is processed in the EU or the wider EEA. The shared Router serves from EU-established providers first, and any machine we rent qualifies only if its data centre sits in that area, with no exception for price or availability. EU Sovereignty Mode goes a step further and restricts processing to sub-processors that are themselves established in the EU, with audit export and support-access controls on the account. Both are set out party by party on our sub-processors page.
2. Dedicated EU Deployment: billed per minute
You pick a GPU class and region, deploy your own vLLM instance, and pay for as long as it runs. Best for custom Hugging Face models, BYOK upstreams, steady high-volume workloads, or when you need full control over the deployment.
The prices below are hourly rates, but you are billed per minute, with no rounding up to the full hour. Stop an instance after six minutes and you pay for six minutes.
GPU pricing follows live EU availability at our providers, so we quote it per deployment rather than from a fixed table. The exact hourly price is shown before you deploy. Ask us for a class we should keep warm for you.
Rekenmodule: wanneer wordt dedicated goedkoper?
Schuif je maandelijkse tokenvolume naar de plek waar jouw werklast zit en kies een modelklasse. De module rekent met de tarieven die hierboven op deze pagina staan, dus met de prijzen waarmee we werkelijk afrekenen.
3. Private single-tenant: on request
Need an isolated runtime with dedicated GPUs per customer, at-rest encryption, and a private network policy? For healthcare, government, legal, finance, and workloads that cannot use shared capacity, we scope and price this per project. The configurations below are typical starting points, not a self-serve product.
| Configuration | VRAM | Indicative / month | Setup (one-off) |
|---|---|---|---|
| 1× L40S | 48 GB | from € 1,200 | € 500 |
| 1× H100 | 80 GB | from € 3,500 | € 1,000 |
| 2× H100 | 160 GB | from € 6,500 | € 1,000 |
| 4× H100 | 320 GB | from € 12,500 | € 1,500 |
Indicative, scoped per project. Talk to us via /contact. Confidential computing (TEE) is on the roadmap; we will not price what we have not yet validated.
BYOK: bring your own API key
You can attach your own OpenAI, Anthropic, Google or Mistral API key to an instance. We forward your traffic to the upstream under your contract with them. BYOK currently carries no platform fee: you only pay your own provider. Useful for hybrid setups that mix EU-hosted open-weights with frontier closed models.
Getting started
- Creating an account is free. No credit card to sign up.
- Pay as you go from a single prepaid credit balance. No subscription, no minimum.
- Top up with iDEAL, card or SEPA, then call the Router or deploy an instance.
Billing
- Currency: EUR. VAT added where applicable; reverse-charge for EU B2B with valid VAT number.
- Method: Stripe (credit card, iDEAL, SEPA direct debit). Invoices auto-issued from your dashboard.
- Credits: top up in advance; balance is consumed by all three modes from a single pool.
- Volume / partner tier: for €> 2 000 / month in tokens or one or more single-tenant deployments, we offer a partner tier with discounts, SLAs, and a dedicated technical contact. Contact info@hostyourai.com.
What's not on the price list
Bespoke procurement, custom contracts, NEN 7510 / BIO audit packages, white-label / reseller arrangements, and confidential-computing deployments are quoted per project. Talk to us via /contact.