Model garden Router · auf Anfrage Dedicated · auf Anfrage

DeepSeek V4 Flash 0731

Dieses Modell läuft als dediziertes Deployment auf großen GPUs und ist nicht standardmäßig im gemeinsamen Playground verfügbar. Kontaktieren Sie uns, wir richten es für Sie ein.

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

deepseek-ai/DeepSeek-V4-Flash-0731 Auf Anfrage
text->text · deepseek-ai · EU-hosted
304B
Parameter
1,049K
Kontextfenster
367GB
Minimaler VRAM
POST /api/v1/chat/completions Auf Anfrage

Spezifikationen

Parameter 304B
Kontextfenster 1,048,576 tokens
Minimaler VRAM 367 GB
Architektur DeepseekV4ForCausalLM (vLLM)
Lizenz mit
Modalität text->text
Veröffentlicht July 2026
Anbieter deepseek-ai ↗

Preise

Gemeinsamer Router · pro Token
Auf Anfrage
Nicht über den gemeinsamen Router verfügbar. Preis auf Anfrage als dediziertes GPU-Deployment.
Dediziertes GPU · pro Stunde
Auf Anfrage
Dediziertes Deployment, ab 367 GB VRAM. Abrechnung nach GPU-Stunden.

Gemeinsamer EU-Router, Pay-per-Token, Scale-to-Zero. Dedizierte GPU-Deployments werden stundenweise abgerechnet, siehe Preise.

Direkt aufrufen

Drop-in-Ersatz für OpenAI: Ändern Sie nur die Base-URL und den API-Key. Auch das Anthropic-Format (/v1/messages) wird unterstützt.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-V4-Flash-0731",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Häufig gestellte Fragen

Kann ich DeepSeek V4 Flash 0731 in der EU betreiben?

Ja. HostYourAI betreibt DeepSeek V4 Flash 0731 auf GPUs in europäischen Rechenzentren über vLLM. Prompts und Outputs verlassen die EU nicht und es ist kein US-Cloud-Anbieter in der Kette.

Ist das Hosting von DeepSeek V4 Flash 0731 DSGVO-konform?

Ja. Die gesamte Verarbeitung findet innerhalb der EU statt, ein Auftragsverarbeitungsvertrag (AVV) ist verfügbar und die Liste der Subprozessoren ist öffentlich. Open-Source-Gewichte bedeuten außerdem: kein Training mit Ihren Daten.

Was kostet DeepSeek V4 Flash 0731?

DeepSeek V4 Flash 0731 ist auf Anfrage verfügbar: über den gemeinsamen EU-Router nach kurzer Abstimmung oder als dediziertes GPU-Deployment, das pro Stunde abgerechnet wird. Nennen Sie uns Ihren Anwendungsfall, dann richten wir es für Sie ein.

Ist die API OpenAI-kompatibel?

Ja. Sie verwenden die Standard-OpenAI-SDKs mit einer angepassten Base-URL (https://hostyourai.com/api/v1). Auch die Anthropic Messages API wird als Drop-in unterstützt.

Weitere Modelle von DeepSeek

DeepSeek V4 Pro DSpark

Note: DeepSeek-V4-Pro-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec

1650B 1M Kontext Modell ansehen →
DeepSeek V4 Flash DSpark

Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec

165B 1M Kontext Modell ansehen →
DeepSeek V4 Pro

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

1599B 1M Kontext Modell ansehen →
DeepSeek V4 Flash

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

158B 1M Kontext Modell ansehen →
DeepSeek OCR 2

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

3.4B 8K Kontext Modell ansehen →
DeepSeek V3.2

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

685B 164K Kontext Modell ansehen →

Zugang anfragen

DeepSeek V4 Flash 0731 ist noch nicht standardmäßig verfügbar. Hinterlassen Sie Ihre Kontaktdaten, wir kümmern uns um ein dediziertes Deployment.

Zugang anfragen