Model garden Router · warm Dedicated · verfügbar

gemma 4 26B A4B it

Sofort über den EU-Router oder als dediziertes GPU-Deployment. Ihre Daten bleiben in Europa.

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned vari...

google/gemma-4-26B-A4B-it vLLM ready
text+image->text · google · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
27B
Parameter
262K
Kontextfenster
80GB
Minimaler VRAM
POST /api/v1/chat/completions 200 OK

Spezifikationen

Parameter 27B
Kontextfenster 262,144 tokens
Minimaler VRAM 80 GB
Architektur Gemma4ForConditionalGeneration (vLLM)
Lizenz apache-2.0
Modalität text+image->text
Veröffentlicht March 2026
Anbieter google ↗

Preise

Gemeinsamer Router · pro Token
€0.29
Input (pro 1M Tokens)
€0.58
Output (pro 1M Tokens)
Dediziertes GPU · pro Stunde
ab €1,68 pro Stunde
Eigene vLLM-Instanz in europäischer Cloud (80 GB VRAM), stundenweise abgerechnet.

Gemeinsamer EU-Router, Pay-per-Token, Scale-to-Zero. Dedizierte GPU-Deployments werden stundenweise abgerechnet, siehe Preise.

✓ Funktionsfähig verifiziert am 27-07-2026, Antwort in 5655 ms auf unserer EU-Infrastruktur.

Direkt aufrufen

Drop-in-Ersatz für OpenAI: Ändern Sie nur die Base-URL und den API-Key. Auch das Anthropic-Format (/v1/messages) wird unterstützt.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "google/gemma-4-26B-A4B-it",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Häufig gestellte Fragen

Kann ich gemma 4 26B A4B it in der EU betreiben?

Ja. HostYourAI betreibt gemma 4 26B A4B it auf GPUs in europäischen Rechenzentren über vLLM. Prompts und Outputs verlassen die EU nicht und es ist kein US-Cloud-Anbieter in der Kette.

Ist das Hosting von gemma 4 26B A4B it DSGVO-konform?

Ja. Die gesamte Verarbeitung findet innerhalb der EU statt, ein Auftragsverarbeitungsvertrag (AVV) ist verfügbar und die Liste der Subprozessoren ist öffentlich. Open-Source-Gewichte bedeuten außerdem: kein Training mit Ihren Daten.

Was kostet gemma 4 26B A4B it?

Über den gemeinsamen EU-Router zahlen Sie €0.29 pro Million Input-Tokens und €0.58 pro Million Output-Tokens, ohne Fixkosten. Für hohe Volumen oder Isolation können Sie gemma 4 26B A4B it auch als dedizierte GPU-Instanz stundenweise betreiben.

Ist die API OpenAI-kompatibel?

Ja. Sie verwenden die Standard-OpenAI-SDKs mit einer angepassten Base-URL (https://hostyourai.com/api/v1). Auch die Anthropic Messages API wird als Drop-in unterstützt.

Weitere Modelle von Google

gemma 4 31B it qat w4a16 ct

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

34B 262K Kontext Modell ansehen →
gemma 4 26B A4B it qat q4 0 unquantized

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

27B 262K Kontext Modell ansehen →
gemma 4 31B it qat q4 0 unquantized

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

33B 262K Kontext Modell ansehen →
gemma 4 31B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

33B 262K Kontext Modell ansehen →
gemma 4 26B A4B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

27B 262K Kontext Modell ansehen →
gemma 4 31B it

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

31B 262K Kontext Modell ansehen →

Testen Sie gemma 4 26B A4B it kostenlos

Die Kontoerstellung dauert eine Minute. Testen Sie gemma 4 26B A4B it direkt im Playground.

Kostenlos starten