Drop-in OpenAI and Anthropic API, hosted in the EU

Host your AI in Europe.
Secure, compliant, drop-in.

Deploy your own AI models in Europe, on a dedicated GPU or through the shared Router. Point your OpenAI or Anthropic client at one base URL and you are live.

Built and hosted in the EU
Your app
OpenAI · Anthropic
EU Router
one base URL
Qwen3-8B
shared gateway
Loes (NL)
dedicated GPU
Llama-3.3
single-tenant
drop-in warm · EU
Your request stays in the EU from start to finish. The Router sends it to a warm model and streams the answer straight back.

Open models, served from the EU on infrastructure you control

Loes MetaLlama Alibaba CloudQwen DeepSeekDeepSeek Mistral AIMistral GoogleGemma FLUX.1 SDXL Phi-3 vLLMvLLM Hugging FaceHuggingFace European GPU marketplaces
0%
EU-hosted

Your data and your models stay on European GPUs. GDPR-friendly from the ground up.

0+
Verified models, ready to serve

Llama, Qwen, DeepSeek, Mistral, FLUX and many more. Pick one and it is warm within minutes, with no DevOps on your side.

0 SDKs
OpenAI and Anthropic compatible

Point your existing client at the Router and keep your tools. No rewrites, no lock-in.

Product

One platform, from first test to production

Router

One endpoint, every open model

Change only the base URL and keep your OpenAI tooling.

chat.completions
# works with your existing OpenAI client curl https://hostyourai.com/api/v1/chat/completions \ -H "Authorization: Bearer hyai-..." \ -d '{ "model": "qwen3-8b", "messages": [{ "role": "user", ... }], "stream": true }' # or your Anthropic SDK, same Router curl https://hostyourai.com/api/v1/messages \ -H "x-api-key: hyai-..."
Model Garden

Browse, compare, deploy

More than 390 serveable open models with live status.

/models
Qwen3 8BwarmEU
Llama 3.3 70BwarmEU
Mistral SmallEU
DeepSeek R1warming upEU
Gemma 3 27BEU
FLUX.1 schnellwarmEU
Phi 4 MiniEU
Qwen2.5 32BEU
Playground

Test any model instantly

Chat with any model before you write a single line of code.

qwen3-8b · streaming · EU
Summarize our return policy in two sentences.
Returns are processed within 14 days of the request. Items must be returned unused and with the original receipt.
And what about discounted items?
Discounted items can be exchanged within the same period.
Activity

Track every request

Usage, latency and cost per request in your activity log.

Requests this weeklatency · cost
qwen3-8b412 ms
llama-3.3-70b890 ms
mistral-small365 ms
gemma-3-27b508 ms
deepseek-r11240 ms
Platform

Everything you need to ship

From your first request to production traffic, you get every model, every endpoint and every insight your team needs in one place.

EU Inference Router

One API. Every open model.

A shared OpenAI-compatible gateway that routes your requests to open models on European GPUs.

OpenAI-compatible API
Automatic routing to EU GPU instances
Anthropic SDK drop-in
Usage and activity log per request
Optional RAG context injection
Explore the Router
EU Inference Router
Incoming /v1/chat request
Authenticate hyai- API key
Pick the nearest warm instance
vLLM streams the response
if (instance.warm === true)
TrueServe right away
FalseWarm up, then route
qwen3-8b vLLM ready
NVIDIA A100 · 40GB · European GPU marketplace · EU region
VRAM19.2 / 40 GB
GPU usage71%
42 ms
time-to-first-token
128
tokens / sec
62°C
temperature
POST /api/v1/chat/completions200 OK
Dedicated Instances

Your own GPU, your own model.

Deploy LLMs (Llama, Qwen, DeepSeek) and image models (FLUX, SDXL) on dedicated GPUs with vLLM. Ready within minutes.

Any HuggingFace model by ID
vLLM on European GPU marketplaces
Auto-generated setup scripts
Warm when you are around, idle when you are not
Private, encrypted upstream keys
Built-in readiness probes
Deploy an instance
Model Garden

Browse, compare, deploy.

A curated catalog of serveable open models with live warm, EU and warming-up status. You always know what is ready.

Curated, serveable catalog
Live warm / EU / warming-up status
A landing page for every model
Verified before you build
Playground for instant testing
Image and chat models in one place
Explore the Model Garden
Model Garden
Chat models
Image models
Embeddings
Warm now
Qwen3-8B
Llama-3.2-1B
Gemma-2-9B
Recently added
DeepSeek-V3
FLUX.1-schnell
Serveable
SDXL-Turbo
Phi-3-mini
How it works

From zero to a warm endpoint in minutes

No infrastructure to manage. Pick a model, get an OpenAI-compatible URL, ship.

01Pick a model

Choose from the Model Garden or paste a HuggingFace ID

Set the VRAM and pick a European GPU. More than 390 verified open models are ready to go.

02Get your endpoint

We deploy vLLM and run readiness probes

You get a warm OpenAI- and Anthropic-compatible URL plus an API key. No DevOps on your side.

03Route and ship

Point your client at the Router

It routes automatically to a warm instance and speaks both the OpenAI and the Anthropic API. Only the base URL changes.

04Track and scale

Every request logged, GPUs idle when inactive

You see usage, latency and cost per request. Instances idle automatically when nobody is online, so you only pay for what you run.

Features

Everything you need for AI

From model hosting to a customer-facing API, built for developers and companies that want to run their AI inside the EU.

OpenAI-compatible endpoint

A drop-in replacement for the OpenAI SDK. Only the base URL changes.

POST /api/v1/chat/completions

Anthropic Messages drop-in

Your Anthropic SDK works out of the box too, including x-api-key auth.

POST /api/v1/messages

Streaming responses

Token by token back to your app, just like you are used to.

stream: true

Your own API keys

Create keys with the hyai- prefix and manage them per project.

Authorization: Bearer hyai-...

390+ model catalog

Llama, Qwen, Mistral, DeepSeek and Gemma, curated and verified.

hostyourai.com/models

Warm-pool

A warm model is always standing by, so your first request never waits.

status: warm

Scale to zero

Instances idle when inactive, so you only pay for what you run.

idle when inactive

Per-request activity log

Usage, latency and cost for every request, live in your dashboard.

GET /router/traces

RAG context injection

Optionally link a knowledge base and the Router injects context automatically.

knowledge_base_id: 42

Any HuggingFace model

Paste a repo ID and deploy your own or fine-tuned model on an EU GPU.

org/model-name

Readiness probes

Every deploy is automatically health-checked before you put traffic on it.

GET /health

Embedding models

Serve embeddings from the same stack, for search and RAG.

POST /api/v1/embeddings

Image models

FLUX and SDXL on dedicated GPUs, alongside your chat models.

FLUX · SDXL

One prepaid balance

Top up with iDEAL, card or SEPA. No subscription, no minimum.

pay-as-you-go

Playground

Try any model in the browser before you put it in your app.

?model=qwen3-8b
Who it's for

Built for teams that cannot send data away

When a US cloud is not an option, HostYourAI gives you the same developer experience on European infrastructure.

Classical government building with columns Public sector

Government & public sector

Citizen data that must legally stay in the EU, fully auditable.

Healthcare

Patient data stays within the EU, on infrastructure with a DPA and a public subprocessor list.

Regulated enterprise

Finance, healthcare and legal teams under GDPR, DORA and the AI Act.

EU SaaS & scale-ups

Ship AI features your customers can trust, without a US sub-processor.

Agencies & integrators

Deliver private AI for clients on infrastructure you can stand behind.

Finance & legal

Open models you can audit, instead of a closed black box.

Row of European flags in front of an EU building Sovereignty

No US cloud in the chain

All inference runs on European GPUs. No US CLOUD Act exposure, no data leaving the EU.

Security

Private from the ground up

HostYourAI keeps your models, prompts and data on European GPUs. Built for teams that care about compliance, reliability and real control.

GDPR-compliant DPA available AES-256 at rest TLS in transit 99.9% uptime SLA European data centers

EU data residency

Prompts and outputs never leave the EU. All inference runs in European data centers.

Encryption everywhere

AES-256 for data at rest and TLS for all traffic in transit.

No training on customer data

Your prompts and outputs are never used to train models.

DPA and subprocessors

A data processing agreement is available and the subprocessor list is public.

99.9% uptime SLA

With a public status page, so you can always see what is warm.

Guides

Get hands-on with the Router

Practical steps to migrate, deploy and build on EU GPUs.

Host. Route. Ship.

Pay as you go and stop whenever you want. No subscription, no minimum.