Pricing

OpenAI-compatible inference with clear tenancy choices: shared, dedicated, or private.

HostYourAI offers three execution modes side by side on a single account and credit balance. Start with EU Hosted inference, move to dedicated capacity when the workload or compliance profile requires it.

1. EU Hosted Gateway: pay per token

One OpenAI-compatible API key, model catalog, hyai/auto, and scale-to-zero shared capacity. Best for SaaS integrations, agencies, agent apps, and experimentation. De tarieven per model (EUR per miljoen tokens), exact de prijzen waarmee je verbruik wordt afgerekend:

Chat- en visionmodellen

Model Context Invoer €/M Cached €/M Uitvoer €/M
Pixtral 12B 125K 0.23 0.23
Qwen3.5 9B 256K 0.17 0.05 0.23
Qwen2.5 14B Instruct AWQ 32K 0.08 0.25
DeepSeek V4 Flash 1M 0.29 0.07 0.35
Mistral Small 3.2 24B 125K 0.17 0.40
Gemma 3 27B 40K 0.29 0.58
DeepSeek V3.2 160K 0.35 0.09 0.58
GPT-OSS 120B 128K 0.16 0.63
Qwen3 235B A22B Instruct 256K 0.21 0.63
Qwen2.5 VL 72B 125K 0.26 0.79
Qwen3 Coder 30B A3B 128K 0.23 0.92
Llama 3.3 70B Instruct 128K 1.04 1.04
MiniMax M2.5 192K 0.35 0.09 1.38
MiniMax M3 200K 0.46 0.12 2.30
DeepSeek R1 0528 160K 0.76 0.20 2.99
Kimi K2.5 256K 0.58 0.15 3.22
GLM 5 198K 1.15 0.29 3.68
Qwen3.5 122B A10B 256K 0.58 0.15 4.03
DeepSeek V4 Pro 1M 2.01 0.51 4.03
Qwen3.5 397B A17B 256K 0.69 4.14
Kimi K2.6 256K 1.15 0.29 4.60
GLM 5.1 198K 1.48 4.66
GLM 5.2 1M 1.73 0.44 5.18
Kimi K2.7 Code 256K 1.44 0.36 5.18
Mistral Medium 3.5 128K 1.73 8.63
Kimi K3 1M 3.45 0.86 17.25

Embeddings

Model Context Invoer €/M tokens
Qwen3 Embedding 8B 32K 0.12
BGE Multilingual Gemma2 8K 0.12

Embeddings hebben geen output-tokens; je betaalt alleen voor de input.

Audio naar tekst

Model Prijs per minuut audio
Whisper Large v3 € 0,0035

Afgerekend per seconde audio, niet per minuut afgerond.

Spraak-, beeld- en videomodellen zijn op aanvraag beschikbaar via /contact.

Modellen zonder gedeelde tokenprijs draaien als dedicated deployment en worden per project geoffreerd via /contact.

EU Hosted means the request is processed in the EU or the wider EEA. The shared Router serves from EU-established providers first, and any machine we rent qualifies only if its data centre sits in that area, with no exception for price or availability. EU Sovereignty Mode goes a step further and restricts processing to sub-processors that are themselves established in the EU, with audit export and support-access controls on the account. Both are set out party by party on our sub-processors page.

2. Dedicated EU Deployment: billed per minute

You pick a GPU class and region, deploy your own vLLM instance, and pay for as long as it runs. Best for custom Hugging Face models, BYOK upstreams, steady high-volume workloads, or when you need full control over the deployment.

The prices below are hourly rates, but you are billed per minute, with no rounding up to the full hour. Stop an instance after six minutes and you pay for six minutes.

GPU pricing follows live EU availability at our providers, so we quote it per deployment rather than from a fixed table. The exact hourly price is shown before you deploy. Ask us for a class we should keep warm for you.

Rekenmodule: wanneer wordt dedicated goedkoper?

Schuif je maandelijkse tokenvolume naar de plek waar jouw werklast zit en kies een modelklasse. De module rekent met de tarieven die hierboven op deze pagina staan, dus met de prijzen waarmee we werkelijk afrekenen.

1,0 miljard tokens per maand
Gedeelde Router goedkoopst
€ 243 per maand
€ 0,17 invoer en € 0,40 uitvoer per miljoen tokens
Dedicated GPU goedkoopst
€ 767 per maand
1x A100 / H100 (80 GB), € 1,05 per uur, 730 uur per maand
Vanaf ongeveer 3,2 miljard tokens per maand is een dedicated GPU goedkoper dan de gedeelde Router voor Mistral Small 3.2 24B.
Bij 1,0 miljard tokens per maand betaal je het minst via de gedeelde Router: € 243 tegen € 767 voor een eigen GPU. Dat scheelt € 524 per maand.
Gerekend met 70 procent invoertokens en 30 procent uitvoertokens, een gangbare verhouding voor chat- en agentverkeer. Zit jouw werklast anders in elkaar, dan schuift het omslagpunt mee.
Dedicated staat altijd aan. Je betaalt de machine per minuut zolang hij draait, ook in het weekend en 's nachts. De gedeelde Router rekent alleen af wat je werkelijk verstuurt en zakt tussendoor terug naar nul.
Dit vergelijkt prijzen, geen capaciteit. Of één machine jouw volume ook werkelijk verwerkt hangt af van je verkeer: gelijktijdigheid, promptlengte en pieken. Bij grote volumes rekenen we dat met je door, zodat je weet hoeveel machines je nodig hebt. /contact
Dedicated geeft je de hele machine. Geen wachtrij achter andere klanten, een voorspelbare latentie en al het GPU-geheugen voor jouw model. Ook bij een gelijk bedrag kan dat de doorslag geven.
Uurprijzen komen uit dezelfde staffel als de tabel hierboven, die dagelijks opnieuw wordt ingelezen bij onze EU-providers. Tokenprijzen komen uit de catalogus waaruit je verbruik wordt afgerekend. Een maand telt hier 730 uur, het jaargemiddelde. Bedragen zijn exclusief btw.

3. Private single-tenant: on request

Need an isolated runtime with dedicated GPUs per customer, at-rest encryption, and a private network policy? For healthcare, government, legal, finance, and workloads that cannot use shared capacity, we scope and price this per project. The configurations below are typical starting points, not a self-serve product.

ConfigurationVRAMIndicative / monthSetup (one-off)
1× L40S48 GBfrom € 1,200€ 500
1× H10080 GBfrom € 3,500€ 1,000
2× H100160 GBfrom € 6,500€ 1,000
4× H100320 GBfrom € 12,500€ 1,500

Indicative, scoped per project. Talk to us via /contact. Confidential computing (TEE) is on the roadmap; we will not price what we have not yet validated.

BYOK: bring your own API key

You can attach your own OpenAI, Anthropic, Google or Mistral API key to an instance. We forward your traffic to the upstream under your contract with them. BYOK currently carries no platform fee: you only pay your own provider. Useful for hybrid setups that mix EU-hosted open-weights with frontier closed models.

Getting started

  • Creating an account is free. No credit card to sign up.
  • Pay as you go from a single prepaid credit balance. No subscription, no minimum.
  • Top up with iDEAL, card or SEPA, then call the Router or deploy an instance.

Billing

  • Currency: EUR. VAT added where applicable; reverse-charge for EU B2B with valid VAT number.
  • Method: Stripe (credit card, iDEAL, SEPA direct debit). Invoices auto-issued from your dashboard.
  • Credits: top up in advance; balance is consumed by all three modes from a single pool.
  • Volume / partner tier: for €> 2 000 / month in tokens or one or more single-tenant deployments, we offer a partner tier with discounts, SLAs, and a dedicated technical contact. Contact info@hostyourai.com.

What's not on the price list

Bespoke procurement, custom contracts, NEN 7510 / BIO audit packages, white-label / reseller arrangements, and confidential-computing deployments are quoted per project. Talk to us via /contact.