Swapping one base URL is all it takes.
You keep your tool and swap the model: Claude Code, Cline and opencode keep working as-is, but you replace the Claude models behind them with open-weight models on European GPUs. No migration, no SDK switch. Pay per token from prepaid credits.
The tool in your terminal stays the same. What changes: the model that answers, and where your code goes.
The router speaks both the Anthropic and the OpenAI API in full. Nothing to rewrite: any tool or SDK that talks to either one works right away.
Before we promise anything on this page we run it ourselves: a full Claude Code session against this router that reads a file through tools and returns the right answer. The example here is that test.
Fair is fair: the US frontier models remain better. This is the alternative for teams that want code and prompts to stay in Europe, with open weights you can inspect yourself.
No black box: you know which model runs, with how much context and which tool support. GET /v1/models shows live what is being served. When a better open coding model appears we verify it and add it to the list.
Your prompts and code are never used to train models. What you send is yours and stays yours.
Inference runs on GPUs in the EU. Your traffic never leaves Europe, not even for logging or analytics.
A DPA is available for teams that need this on paper. Email info@hostyourai.com.
You top up credits from €5 by iDEAL or card and pay per token, priced per model. When it runs out it stops: your agent can never spend more than you set aside. An afternoon with a coding agent typically costs cents to a few euros.
Current token prices are on the pricing page.
No. Claude Code is only the tool in your terminal; you replace the Claude models behind it with open-weight models such as Qwen, hosted by us in the EU. Your traffic no longer goes to Anthropic.
No. Anything that speaks the Anthropic or OpenAI API keeps working with a different base URL and key. Claude Code, Cline and opencode only need three environment variables.
Currently a Qwen3.6 coding model with tools and 32k context, the best open-weight model we have verified. The list grows as better open models appear.
A warm GPU responds quickly and streams token by token. If the model is cold, your first question boots a GPU and takes a few minutes. After that it stays warm while you work.
Per token from your prepaid credit balance. An afternoon with a coding agent typically costs cents to a few euros; every request shows up live in your usage.
No. No subscription; you top up whenever you want and your balance stays. And because the APIs are standard, switching back is just as easy.
Credits from €5, no subscription. Your first session runs within minutes.