Model garden Router · on request Dedicated · available

DeepSeek OCR 2

Runs as your own dedicated GPU deployment, ready to launch from the wizard. Data stays in Europe.

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

deepseek-ai/DeepSeek-OCR-2 vLLM ready
text+image->text · deepseek-ai · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
3.4B
Parameters
8K
Context window
16GB
Minimum VRAM
POST /api/v1/chat/completions On request

Specifications

Parameters 3.4B
Context window 8,192 tokens
Minimum VRAM 16 GB
Architecture DeepseekOCR2ForCausalLM (vLLM)
License apache-2.0
Modality text+image->text
Released January 2026
Publisher deepseek-ai ↗

Pricing

Shared router · per token
On request
Not available on the shared router. Pricing on request as a dedicated GPU deployment.
Dedicated GPU · per hour
from €1,95 per hour
Your own vLLM instance on European cloud (16 GB VRAM), billed hourly.

Shared EU router, pay-per-token, scale-to-zero. Dedicated GPU deployments are billed hourly, see pricing.

✓ Verified working on 27-07-2026, responded in 438 ms on our EU infrastructure.

Call it now

Drop-in replacement for OpenAI: change only the base URL and API key. The Anthropic format (/v1/messages) is supported too.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "deepseek-ai/DeepSeek-OCR-2",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Frequently asked questions

Can I run DeepSeek OCR 2 in the EU?

Yes. HostYourAI runs DeepSeek OCR 2 on GPUs in European datacenters via vLLM. Prompts and outputs never leave the EU and there is no US cloud provider in the chain.

Is hosting DeepSeek OCR 2 GDPR-compliant?

Yes. All processing happens inside the EU, a Data Processing Agreement (DPA) is available and the subprocessor list is public. Open-source weights also mean: no training on your data.

How much does DeepSeek OCR 2 cost?

DeepSeek OCR 2 is available on request: through the shared EU router after a quick chat, or as your own dedicated GPU deployment billed per hour. Tell us your use case and we will set it up for you.

Is the API OpenAI-compatible?

Yes. You use the standard OpenAI SDKs with a custom base URL (https://hostyourai.com/api/v1). The Anthropic Messages API is supported as a drop-in as well.

More models from DeepSeek

DeepSeek V4 Flash 0731

DeepSeek-V4-Flash-0731 is the official release of DeepSeek-V4-Flash, superseding the preview version, with substantially enhanced agentic capabilities. It has the same model structure as DeepSeek-V4-Flash-DSpark, i.e. it comes with a speculative decoding module attached.

304B 1M context View model →
DeepSeek V4 Pro

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

1M context View model →
DeepSeek V4 Flash

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

291B 1M context View model →
DeepSeek V3.2

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

164K context View model →
DeepSeek OCR

torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation

3.3B 8K context View model →
DeepSeek V3.2 Exp

We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.

164K context View model →

Try DeepSeek OCR 2 for free

Creating an account takes a minute. Test DeepSeek OCR 2 straight away in the playground.

Start for free