Model garden Router · on request Dedicated · available

GLM 5 FP8

Runs as your own dedicated GPU deployment, ready to launch from the wizard. Data stays in Europe.

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (3...

zai-org/GLM-5-FP8 On request
text->text · zai-org · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
Parameters
203K
Context window
908GB
Minimum VRAM
POST /api/v1/chat/completions On request

Specifications

Context window 202,752 tokens
Minimum VRAM 908 GB
Architecture GlmMoeDsaForCausalLM (vLLM)
License open-weights
Modality text->text
Released February 2026
Publisher zai-org ↗

Pricing

Shared router · per token
On request
Not available on the shared router. Pricing on request as a dedicated GPU deployment.
Dedicated GPU · per hour
from €51,77 per hour
Your own vLLM instance on European cloud (908 GB VRAM), billed hourly.

Shared EU router, pay-per-token, scale-to-zero. Dedicated GPU deployments are billed hourly, see pricing.

Call it now

Drop-in replacement for OpenAI: change only the base URL and API key. The Anthropic format (/v1/messages) is supported too.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "zai-org/GLM-5-FP8",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Frequently asked questions

Can I run GLM 5 FP8 in the EU?

Yes. HostYourAI runs GLM 5 FP8 on GPUs in European datacenters via vLLM. Prompts and outputs never leave the EU and there is no US cloud provider in the chain.

Is hosting GLM 5 FP8 GDPR-compliant?

Yes. All processing happens inside the EU, a Data Processing Agreement (DPA) is available and the subprocessor list is public. Open-source weights also mean: no training on your data.

How much does GLM 5 FP8 cost?

GLM 5 FP8 is available on request: through the shared EU router after a quick chat, or as your own dedicated GPU deployment billed per hour. Tell us your use case and we will set it up for you.

Is the API OpenAI-compatible?

Yes. You use the standard OpenAI SDKs with a custom base URL (https://hostyourai.com/api/v1). The Anthropic Messages API is supported as a drop-in as well.

More models from Z.AI

GLM 5.3 Flash

We introduce GLM-5.3-Flash, the first natively multimodal model in the GLM-5 series. With 320B total parameters and just 18B active parameters, it outperforms GLM-5.2 across benchmarks and real-world workloads at one-tenth the price, while approaching Claude Opus 4.8 on coding and agentic benchmarks.

321B 1M context View model →
GLM 5.3

GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training. Compared with GLM-5.2, it is much better at complex coding and long-horizon tasks:

753B 1M context View model →
GLM 5.2

We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency - Improved Architecture: We propose IndexShare, which reuses the same indexer across every fou

753B 1M context View model →
GLM 5.1

GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

754B 203K context View model →
GLM 5

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

754B 203K context View model →
GLM 4.7 Flash

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

31B 203K context View model →

Try GLM 5 FP8 for free

Creating an account takes a minute. Test GLM 5 FP8 straight away in the playground.

Start for free