Model garden Router · auf Anfrage Dedicated · verfügbar

Qwen2.5 VL 32B Instruct AWQ

Läuft als eigenes dediziertes GPU-Deployment, direkt über den Wizard startbar. Daten bleiben in Europa.

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introdu...

Qwen/Qwen2.5-VL-32B-Instruct-AWQ vLLM ready
text+image->text · Qwen · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
33B
Parameter
128K
Kontextfenster
160GB
Minimaler VRAM
POST /api/v1/chat/completions Auf Anfrage

Spezifikationen

Parameter 33B
Kontextfenster 128,000 tokens
Minimaler VRAM 160 GB
Architektur Qwen2_5_VLForConditionalGeneration (vLLM)
Lizenz apache-2.0
Modalität text+image->text
Veröffentlicht March 2025
Anbieter Qwen ↗

Preise

Gemeinsamer Router · pro Token
Auf Anfrage
Nicht über den gemeinsamen Router verfügbar. Preis auf Anfrage als dediziertes GPU-Deployment.
Dediziertes GPU · pro Stunde
ab €6,97 pro Stunde
Eigene vLLM-Instanz in europäischer Cloud (160 GB VRAM), stundenweise abgerechnet.

Gemeinsamer EU-Router, Pay-per-Token, Scale-to-Zero. Dedizierte GPU-Deployments werden stundenweise abgerechnet, siehe Preise.

✓ Funktionsfähig verifiziert am 20-07-2026, Antwort in 126 ms auf unserer EU-Infrastruktur.

Direkt aufrufen

Drop-in-Ersatz für OpenAI: Ändern Sie nur die Base-URL und den API-Key. Auch das Anthropic-Format (/v1/messages) wird unterstützt.

curl https://hostyourai.com/api/v1/chat/completions \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen2.5-VL-32B-Instruct-AWQ",
    "messages": [{"role": "user", "content": "Hello!"}]
  }'

Häufig gestellte Fragen

Kann ich Qwen2.5 VL 32B Instruct AWQ in der EU betreiben?

Ja. HostYourAI betreibt Qwen2.5 VL 32B Instruct AWQ auf GPUs in europäischen Rechenzentren über vLLM. Prompts und Outputs verlassen die EU nicht und es ist kein US-Cloud-Anbieter in der Kette.

Ist das Hosting von Qwen2.5 VL 32B Instruct AWQ DSGVO-konform?

Ja. Die gesamte Verarbeitung findet innerhalb der EU statt, ein Auftragsverarbeitungsvertrag (AVV) ist verfügbar und die Liste der Subprozessoren ist öffentlich. Open-Source-Gewichte bedeuten außerdem: kein Training mit Ihren Daten.

Was kostet Qwen2.5 VL 32B Instruct AWQ?

Qwen2.5 VL 32B Instruct AWQ braucht mehrere GPUs gleichzeitig und läuft deshalb als dediziertes Deployment, das nach GPU-Stunden abgerechnet wird, nicht pro Token. Nennen Sie uns Ihr Volumen, dann rechnen wir es für Sie durch.

Ist die API OpenAI-kompatibel?

Ja. Sie verwenden die Standard-OpenAI-SDKs mit einer angepassten Base-URL (https://hostyourai.com/api/v1). Auch die Anthropic Messages API wird als Drop-in unterstützt.

Weitere Modelle von Qwen

Qwen3.8 27B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc.

28B 262K Kontext Modell ansehen →
Qwen3 ASR 0.6B hf

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

0.8B 66K Kontext Modell ansehen →
Qwen3 ASR 1.7B hf

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

2B 66K Kontext Modell ansehen →
Qwen3.6 27B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

28B 262K Kontext Modell ansehen →
Qwen3.6 35B A3B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

36B 262K Kontext Modell ansehen →
Qwen3.5 0.8B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intended use cases are prototyping, task-specific fine-tuning, and other research or development purposes.

0.9B 262K Kontext Modell ansehen →

Testen Sie Qwen2.5 VL 32B Instruct AWQ kostenlos

Die Kontoerstellung dauert eine Minute. Testen Sie Qwen2.5 VL 32B Instruct AWQ direkt im Playground.

Kostenlos starten