Model garden Router · warm Dedicated · beschikbaar

gemma 4 26B A4B it

Direct via de EU-router of als dedicated GPU-deployment. Data blijft in Europa.

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned vari...

google/gemma-4-26B-A4B-it vLLM ready
text+image->text · google · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
26B
Parameters
262K
Contextvenster
80GB
Minimale VRAM
POST /api/v1/chat/completions 200 OK

Specificaties

Parameters 26B
Contextvenster 262,144 tokens
Minimale VRAM 80 GB
Architectuur Gemma4ForConditionalGeneration (vLLM)
Licentie apache-2.0
Modaliteit text+image->text
Uitgebracht March 2026
Uitgever google ↗

Prijzen

Gedeelde router · per token
€0.29
Input (per 1M tokens)
€0.58
Output (per 1M tokens)
Dedicated GPU · per uur
vanaf €2,42 per uur
Eigen vLLM-instance op Europese cloud (80 GB VRAM), per uur afgerekend.

Gedeelde EU-router, pay-per-token, scale-to-zero. Dedicated GPU-deployments worden per uur afgerekend, zie prijzen.

✓ Werkend geverifieerd op 27-07-2026, respons in 5655 ms op onze EU-infrastructuur.

Bench-index

Bench is onze kwaliteitsindex voor elk model, 0 tot 100. Hij steunt op openbare metingen (Epoch AI met zo’n 40 benchmarks, LiveBench, LMArena en meer) en wordt elke ochtend ververst. Waar niemand meet, staat een schatting met bandbreedte.

64,8 / 100 · Gemeten

Gemeten door Epoch AI over 6 benchmarks, met bandbreedte.

Bandbreedte 60,3 tot 66,9 · Plek 101 van 268 bij Epoch AI

Bronnen: Epoch AI ↗ · LiveBench ↗ · LMArena ↗ · Dagelijks bijgewerkt, laatst op 27-09-2026

Alle benchmarks (6)
OTIS Mock AIME 2024-2025 82,2
GPQA diamond 64,3
DTBench 58,2
WeirdML 35,2
LMCA 34,9
Chess Puzzles 1,1

Direct aanroepen

Drop-in vervanger voor OpenAI: wijzig alleen de base-URL en de API-key. Ook het Anthropic-formaat (/v1/messages) wordt ondersteund.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-26B-A4B-it",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Veelgestelde vragen

Kan ik gemma 4 26B A4B it in de EU draaien?

Ja. HostYourAI draait gemma 4 26B A4B it op GPU's in Europese datacenters via vLLM. Prompts en outputs verlaten de EU niet en er is geen Amerikaanse cloudprovider in de keten.

Is gemma 4 26B A4B it hosten AVG/GDPR-compliant?

Ja. Alle verwerking vindt plaats binnen de EU, er is een verwerkersovereenkomst (DPA) beschikbaar en de subprocessor-lijst is openbaar. Open-source gewichten betekenen ook: geen training op jouw data.

Wat kost gemma 4 26B A4B it?

Via de gedeelde EU-router betaal je €0.29 per miljoen input-tokens en €0.58 per miljoen output-tokens, zonder vaste kosten. Voor hoge volumes of isolatie kun je gemma 4 26B A4B it ook als dedicated GPU-instance per uur draaien.

Is de API compatibel met OpenAI?

Ja. Je gebruikt de standaard OpenAI-SDK's met een aangepaste base-URL (https://hostyourai.com/api/v1). Ook de Anthropic Messages API wordt ondersteund als drop-in.

Andere modellen van Google

diffusiongemma 26B A4B it

DiffusionGemma is a generative model built by Google DeepMind. Based on the 26B A4B Mixture-of-Experts (MoE) Gemma 4 architecture, DiffusionGemma generates tokens using discrete diffusion. This open-weights model is multimodal, handling text, image, and video inputs to generate text output.

26B 262K context Bekijk model →
gemma 4 31B it qat w4a16 ct

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

34B 262K context Bekijk model →
gemma 4 26B A4B it qat q4 0 unquantized

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

27B 262K context Bekijk model →
gemma 4 31B it qat q4 0 unquantized

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

33B 262K context Bekijk model →
gemma 4 31B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

33B 262K context Bekijk model →
gemma 4 26B A4B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

27B 262K context Bekijk model →

Probeer gemma 4 26B A4B it gratis

Account aanmaken duurt een minuut. Test gemma 4 26B A4B it direct in de playground.

Start gratis