Model garden Router · op aanvraag Dedicated · beschikbaar

Qwen3 Embedding 0.6B

Dit model draait als eigen dedicated GPU-deployment, direct te starten via de wizard. Data blijft in Europa.

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking mod...

Qwen/Qwen3-Embedding-0.6B Op aanvraag
text->embedding · Qwen · sovereign EU
Runs on EU infrastructure operated by European companies; US marketplace capacity is never part of this chain. Full chain
0.6B
Parameters
33K
Contextvenster
8GB
Minimale VRAM
POST /api/v1/embeddings Op aanvraag

Specificaties

Parameters 0.6B
Contextvenster 32,768 tokens
Minimale VRAM 8 GB
Architectuur Qwen3ForCausalLM (vLLM)
Licentie apache-2.0
Modaliteit text->embedding
Uitgebracht June 2025
Uitgever Qwen ↗

Prijzen

Gedeelde router · per token
Op aanvraag
Niet beschikbaar op de gedeelde router. Prijs op aanvraag als dedicated GPU-deployment.
Dedicated GPU · per uur
vanaf €1,95 per uur
Eigen vLLM-instance op Europese cloud (8 GB VRAM), per uur afgerekend.

Gedeelde EU-router, pay-per-token, scale-to-zero. Dedicated GPU-deployments worden per uur afgerekend, zie prijzen.

Direct aanroepen

Drop-in vervanger voor OpenAI: wijzig alleen de base-URL en de API-key. Ook het Anthropic-formaat (/v1/messages) wordt ondersteund.

curl https://hostyourai.com/api/v1/embeddings \
  -H "Authorization: Bearer hyai-..." \
  -H "Content-Type: application/json" \
  -d '{
    "model": "Qwen/Qwen3-Embedding-0.6B",
    "input": "The quick brown fox"
  }'

Veelgestelde vragen

Kan ik Qwen3 Embedding 0.6B in de EU draaien?

Ja. HostYourAI draait Qwen3 Embedding 0.6B op GPU's in Europese datacenters via vLLM. Prompts en outputs verlaten de EU niet en er is geen Amerikaanse cloudprovider in de keten.

Is Qwen3 Embedding 0.6B hosten AVG/GDPR-compliant?

Ja. Alle verwerking vindt plaats binnen de EU, er is een verwerkersovereenkomst (DPA) beschikbaar en de subprocessor-lijst is openbaar. Open-source gewichten betekenen ook: geen training op jouw data.

Wat kost Qwen3 Embedding 0.6B?

Qwen3 Embedding 0.6B heeft meerdere GPU's tegelijk nodig en draait daarom als dedicated deployment. Je betaalt dan per GPU-uur en niet per token. Vertel ons je volume, dan rekenen we het voor je door.

Is de API compatibel met OpenAI?

Ja. Je gebruikt de standaard OpenAI-SDK's met een aangepaste base-URL (https://hostyourai.com/api/v1). Ook de Anthropic Messages API wordt ondersteund als drop-in.

Andere modellen van Qwen

Qwen3 ASR 0.6B hf

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

0.8B 66K context Bekijk model →
Qwen3 ASR 1.7B hf

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

2B 66K context Bekijk model →
Qwen3.6 27B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

28B 262K context Bekijk model →
Qwen3.6 35B A3B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

36B 262K context Bekijk model →
Qwen3.5 0.8B Base

[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger

0.9B 262K context Bekijk model →
Qwen3.5 0.8B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. In light of its parameter scale, the intended use cases are prototyping, task-specific fine-tuning, and other research or development purposes.

0.9B 262K context Bekijk model →

Probeer Qwen3 Embedding 0.6B gratis

Account aanmaken duurt een minuut. Je API-key werkt direct met Qwen3 Embedding 0.6B.

Start gratis