Model garden

Modellkatalog

715 Open-Source-Modelle, gehostet auf GPUs in der EU. Ein OpenAI-kompatibler API-Key, Scale-to-Zero oder dediziert.

715 Modelle

Kimi K3
Moonshotai

Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.

1M Kontext Router €3.17 pro 1M Input GPU ab €90,00/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
DeepSeek V4 Flash
DeepSeek · 291B

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

1M Kontext Router €0.29 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3 Embedding 8B
Qwen

Qwen3 Embedding 8B ist ein quelloffenes Sprachmodell von Qwen mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.12 pro 1M Input GPU ab €1,68/Std. Embeddings In der EU gehostet Router · warm Dedicated · verfügbar
Llama 3.3 70B
Meta · 70B

Llama 3.3 70B ist ein quelloffenes Sprachmodell von Meta mit 70B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.14 pro 1M Input GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 9B
Qwen · 9.7B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

262K Kontext Router €0.17 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
Llama 3.3 70B Instruct
Llama-3.3-70b-instruct

Llama 3.3 70B Instruct ist ein quelloffenes Sprachmodell von Llama-3.3-70b-instruct mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.14 pro 1M Input GPU ab €5,09/Std. In der EU gehostet Router · warm Dedicated · verfügbar
GLM 5.2
Z.AI · 753B

We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency - Improved Architecture: We propose IndexShare, which reuses the same indexer across every fou

1M Kontext Router €1.73 pro 1M Input In der EU gehostet Router · warm Dedicated · auf Anfrage
Qwen3 Coder 30B A3B
Qwen3-coder-30b-a3b-instruct

Qwen3 Coder 30B A3B ist ein quelloffenes Sprachmodell von Qwen3-coder-30b-a3b-instruct mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.23 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3 235B A22B Instruct
Qwen3-235b-a22b-instruct-2507

Qwen3 235B A22B Instruct ist ein quelloffenes Sprachmodell von Qwen3-235b-a22b-instruct-2507 mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.08 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · warm Dedicated · verfügbar
BGE Multilingual Gemma2
BAAI

BGE Multilingual Gemma2 ist ein quelloffenes Sprachmodell von BAAI mit einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

8K Kontext Router €0.12 pro 1M Input GPU ab €1,68/Std. Embeddings In der EU gehostet Router · warm Dedicated · verfügbar
Loes Large
Qwen

Loes Large World: Qwen3.5-27B dense base met volledig-Europese SFT-adapter (LoRA), geserveerd via vLLM --enable-lora + qwen3 reasoning-parser, CUDA-graphs. Topmodel-spoor voor chat.loes.ai.

33K Kontext Router €0.25 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
DeepSeek V4 Pro
DeepSeek

We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.

1M Kontext Router €2.01 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Kimi K2.6
Moonshotai · 1027B

Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.

262K Kontext Router €1.15 pro 1M Input Vision In der EU gehostet Router · warm Dedicated · auf Anfrage
Qwen2.5 32B
Qwen · 32B

Qwen2.5 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Mistral Medium 3.5
Mistral-medium-3.5-128b

Mistral Medium 3.5 ist ein quelloffenes Sprachmodell von Mistral-medium-3.5-128b mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €1.73 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Mistral Nemo 12B
Mistral · 12B

Mistral Nemo 12B ist ein quelloffenes Sprachmodell von Mistral mit 12B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.10 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Mistral Small 3.2 24B
Mistral-small-3.2-24b-instruct-2506

Mistral Small 3.2 24B ist ein quelloffenes Sprachmodell von Mistral-small-3.2-24b-instruct-2506 mit einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext Router €0.17 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
GPT-OSS 120B
Gpt-oss-120b

GPT-OSS 120B ist ein quelloffenes Sprachmodell von Gpt-oss-120b mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.16 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Gemma 4 26B A4B
Gemma-4-26b-a4b-it

Gemma 4 26B A4B ist ein quelloffenes Sprachmodell von Gemma-4-26b-a4b-it mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.29 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Gemma 3 27B
Gemma-3-27b-it

Gemma 3 27B ist ein quelloffenes Sprachmodell von Gemma-3-27b-it mit einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.10 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Loes World
HostYourAI

Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · warm Dedicated · verfügbar
Qwen3.5 122B A10B
Qwen · 125B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

262K Kontext Router €0.58 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
Qwen2.5 VL 72B
Qwen · 72B

Qwen2.5 VL 72B ist ein quelloffenes Sprachmodell von Qwen mit 72B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext Router €0.26 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3 VL 8B
Qwen · 8B

Qwen3 VL 8B ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Holo2 30B A3B
Holo2-30b-a3b

Holo2 30B A3B ist ein quelloffenes Sprachmodell von Holo2-30b-a3b mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.35 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · verfügbar Dedicated · verfügbar
Kimi K2.7 Code
Moonshotai · 1027B

Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.

262K Kontext Router €1.44 pro 1M Input Vision In der EU gehostet Router · warm Dedicated · auf Anfrage
Qwen3.6 35B A3B
Qwen · 35B

Qwen3.6 35B A3B ist ein quelloffenes Sprachmodell von Qwen mit 35B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Pixtral 12B
Pixtral-12b-2409

Pixtral 12B ist ein quelloffenes Sprachmodell von Pixtral-12b-2409 mit einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext Router €0.23 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
DeepSeek R1 0528
DeepSeek

The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro.

164K Kontext Router €0.76 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen2.5 VL 7B
Qwen · 7B

Qwen2.5 VL 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 5.1
Z.AI · 754B

GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).

203K Kontext Router €1.48 pro 1M Input In der EU gehostet Router · warm Dedicated · auf Anfrage
MiniMax M3
MiniMaxAI

MiniMax M3 ist ein quelloffenes Sprachmodell von MiniMaxAI mit einem Kontextfenster von 205K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

205K Kontext Router €0.46 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3.5 397B A17B
Qwen3.5-397b-a17b

Qwen3.5 397B A17B ist ein quelloffenes Sprachmodell von Qwen3.5-397b-a17b mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.69 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Whisper Large v3
OpenAI

Whisper Large v3 ist ein quelloffenes Sprachmodell von OpenAI, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €57.50 pro 1M Input GPU ab €1,68/Std. Transkription In der EU gehostet Router · warm Dedicated · verfügbar
MiniMax M2.5
MiniMaxAI

MiniMax M2.5 ist ein quelloffenes Sprachmodell von MiniMaxAI mit einem Kontextfenster von 197K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

197K Kontext Router €0.35 pro 1M Input In der EU gehostet Router · warm Dedicated · auf Anfrage
DeepSeek V3.2
DeepSeek

We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:

164K Kontext Router €0.35 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen 3 0.6B
Qwen · 0.6B

Qwen 3 0.6B ist ein quelloffenes Sprachmodell von Qwen mit 0.6B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Loes Large EU (EuroLLM-22B-2512)
HostYourAI · 22B

Sovereign EU model fine-tuned by HostYourAI on loes-large-v1.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Devstral 2 123B
Devstral-2-123b-instruct-2512

Devstral 2 123B ist ein quelloffenes Sprachmodell von Devstral-2-123b-instruct-2512 mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.46 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · verfügbar Dedicated · verfügbar
GLM 5
Z.AI · 754B

We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.

203K Kontext Router €1.15 pro 1M Input In der EU gehostet Router · warm Dedicated · auf Anfrage
Kimi K2.5
Moonshotai · 1027B

Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.

262K Kontext Router €0.58 pro 1M Input Vision In der EU gehostet Router · warm Dedicated · auf Anfrage
Qwen 3 Coder 30B-A3B (MoE)
Qwen · 30B

Qwen 3 Coder 30B-A3B (MoE) ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.23 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.6 27B
Qwen · 27B

Qwen3.6 27B ist ein quelloffenes Sprachmodell von Qwen mit 27B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EuroLLM 22B Instruct 2512
Utter-project · 23B

This is the model card for EuroLLM-22B-Instruct. You can also check the pre-trained version: EuroLLM-22B-2515.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Mistral 7B Instruct v0.2
Mistral · 7.2B

py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill 14B
DeepSeek · 14B

DeepSeek R1 Distill 14B ist ein quelloffenes Sprachmodell von DeepSeek mit 14B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 32B
Qwen · 32B

Qwen3 VL 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 1B
Meta · 1B

Llama 3.2 1B ist ein quelloffenes Sprachmodell von Meta mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Loes (EuroLLM-22B)
HostYourAI · 22B

Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · verfügbar Dedicated · verfügbar
Voxtral Small 24B
Voxtral-small-24b-2507

Voxtral Small 24B ist ein quelloffenes Sprachmodell von Voxtral-small-24b-2507 mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.17 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · verfügbar Dedicated · verfügbar
DeepSeek R1 0528 Qwen3 8B
DeepSeek · 8B

DeepSeek R1 0528 Qwen3 8B ist ein quelloffenes Sprachmodell von DeepSeek mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.8 27B FP8
Qwen · 28B

[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.

262K Kontext GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V4 Pro 0813
DeepSeek · 1650B

DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.

1M Kontext In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
NVIDIA Nemotron 3.5 Lightning 30B A3B NVFP4
NVIDIA · 18B

The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of May 2026.

1M Kontext GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AREX Base
BAAI · 123B

AREX is a family of deep research agents developed by the Beijing Academy of Artificial Intelligence (BAAI). It is designed for long-horizon tasks in which an agent must search across sources, assemble candidate answers, verify multiple constraints, and revise its research plan when the available evidence is incomplete.

262K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AREX Turbo
BAAI · 4.5B

AREX is a family of deep research agents developed by the Beijing Academy of Artificial Intelligence (BAAI). It is designed for long-horizon tasks in which an agent must search across sources, assemble candidate answers, verify multiple constraints, and revise its research plan when the available evidence is incomplete.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Fara1.5 27B
Microsoft · 27B

Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Fara1.5 4B
Microsoft · 4.5B

Fara1.5-4B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Ising Calibration 1.5 31B NVFP4
NVIDIA · 31B

GOVERNING TERMS: Use of this model is governed by the OpenMDW License Agreement, version 1.1. ADDITIONAL INFORMATION: Apache License, Version 2.0.

262K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Ising Calibration 1.5 31B BF16
NVIDIA · 31B

GOVERNING TERMS: Use of this model is governed by the OpenMDW License Agreement, version 1.1. ADDITIONAL INFORMATION: Apache License, Version 2.0.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 7b fc 2607
BSC-LT · 7.8B

salamandra 7b fc 2607 ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
ALIA 40b fc 2607
BSC-LT · 40B

ALIA 40b fc 2607 ist ein quelloffenes Sprachmodell von BSC-LT mit 40B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Cosmos3 Super Text2Image 4Step
NVIDIA · 64B

Cosmos3-Super-Text2Image-4Step is a 4-step distilled version of the base Cosmos3-Super-Text2Image model. Given a text prompt, it generates a high-fidelity image.

262K Kontext GPU ab €5,09/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen Image Flash
NVIDIA · 20B

The NVIDIA Qwen-Image-Flash model generates images from text prompts using a four-step, DMD2-distilled version of Qwen/Qwen-Image. The distillation used DMD2 from NVIDIA FastGen, NVIDIA Model Optimizer, and NVIDIA AutoModel while retaining the base model architecture. The packaged scheduler is configured for the four-step, shift-3 trajectory.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 7b instruct 2606
BSC-LT · 7.8B

[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.

164K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
ALIA 40b instruct 2606
BSC-LT · 40B

[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.

164K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
GELab Zero 4B preview Sico Evolution
Microsoft · 4.4B

GELab Zero 4B preview Sico Evolution ist ein multimodales Sprachmodell von Microsoft mit 4.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
ALIA 40b fc 2606
BSC-LT · 40B

[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.

164K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Qwen3 ASR 0.6B hf
Qwen · 0.8B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

66K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 ASR 1.7B hf
Qwen · 2B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.

66K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron Labs 3 Puzzle 75B A9B NVFP4
NVIDIA · 45B

The model employs a hybrid MoE architecture with interleaved Mamba, MoE, and Attention layers. Like Nemotron-3-Super, it supports Multi-Token Prediction (MTP) for faster text generation. Compared to its parent, Puzzle-75B-A9B reduces the model from 120.7B total / 12.8B active parameters to 75.3B total / 9.3B active parameters.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B Long Instruct
LumiOpen · 8B

Poro 2 Long Instruct is an instruction-following chatbot model with extended context support, created through supervised fine-tuning (SFT) of the Poro 2 Long Base model followed by merging the SFT checkpoint back with the base model to preserve long-context performance. This model is designed for conversational AI applications and instruction following in both Finnish and English, with support for context lengths up to 128K tokens. It was trained on a carefully curated mix of English and Finnish instruction data.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 4 31B it qat w4a16 ct
Google · 34B

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

262K Kontext GPU ab €5,09/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Ultra 550B A55B NVFP4
NVIDIA · 550B

For more details on how to deploy and use the model - see the Quick Start Guide below!

262K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Ultra 550B A55B BF16
NVIDIA · 550B

For more details on how to deploy and use the model - see the Quick Start Guide below!

262K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos3 Super Text2Image
NVIDIA · 65B

Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning.

262K Kontext GPU ab €5,09/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B Instruct vLLM NVFP4A16
Swiss-ai · 1.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 4B Instruct MLX INT6
Swiss-ai · 0.8B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 4B Instruct MLX INT4
Swiss-ai · 0.6B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 4B Instruct MLX INT3
Swiss-ai · 0.6B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B Instruct MLX INT6
Swiss-ai · 0.3B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B Instruct MLX INT4
Swiss-ai · 0.3B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B Instruct MLX INT3
Swiss-ai · 0.2B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 0.5B Instruct MLX INT6
Swiss-ai · 0.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 0.5B Instruct MLX INT4
Swiss-ai · 0.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 0.5B Instruct MLX INT3
Swiss-ai · 0.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
ALIA 40b fc 2605
BSC-LT · 40B

[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.

164K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
AnyFlow FAR Wan2.1 14B Diffusers
NVIDIA · 14B

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AnyFlow FAR Wan2.1 1.3B Diffusers
NVIDIA · 1.4B

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AnyFlow Wan2.1 T2V 14B Diffusers
NVIDIA · 14B

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AnyFlow Wan2.1 T2V 1.3B Diffusers
NVIDIA · 1.4B

AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
MagenticBrain
Microsoft · 15B

MagenticBrain is a 14B-parameter orchestration model from Microsoft Research AI Frontiers. It plans multi-step tasks, calls declared tools, and coordinates sub-agents. It does not execute actions itself — every real-world side effect happens inside a host harness.

41K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Fara1.5 9B
Microsoft · 9.4B

Fara1.5-9B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
ALIA 40b instruct 2605
BSC-LT · 40B

[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.

164K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B Long Math Reasoning RL Preview
LumiOpen · 8B

Poro 2 8B Math Reasoning RL Preview is a specialized model focused on mathematical reasoning and problem-solving. This preview model was created through reinforcement learning (RL) on top of the Math Reasoning SFT checkpoint. This model excels at mathematical reasoning tasks but is not optimized for general conversational use or other domains.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B Instruct
Swiss-ai · 1.5B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 4B Instruct
Swiss-ai · 3.8B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 0.5B Instruct
Swiss-ai · 0.6B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Cosmos Reason2 32B
NVIDIA · 33B

NVIDIA Cosmos Reason 2 is an open, customizable, 32B-parameter reasoning vision language model (VLM) for physical AI and robotics that enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.

262K Kontext GPU ab €5,09/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 4 26B A4B it qat q4 0 unquantized
Google · 27B

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 4 31B it qat q4 0 unquantized
Google · 33B

[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 311m multilingual r2
Ibm-granite · 0.3B

Model Summary: Granite-Embedding-311M-Multilingual-R2 is a 311M parameter dense embedding model from the Granite Embeddings collection for high-quality multilingual text embeddings. It produces 768-dimensional vectors with a context length of up to 32,768 tokens. The model supports 200+ languages (based on the multilingual pretraining corpus of the underlying encoder), with enhanced support for 52 languages and programming code that receive explicit retrieval-pair and cross-lingual training. All training data uses permissive, enterprise-friendly licenses, plus IBM-collected and IBM-generated d

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 97m multilingual r2
Ibm-granite · 0.1B

Model Summary: Granite-Embedding-97M-Multilingual-R2 is a 97M parameter dense embedding model from the Granite Embeddings collection for high-quality multilingual text embeddings at minimal compute cost. It produces 384-dimensional vectors with a context length of up to 32,768 tokens. The model supports 200+ languages (based on the multilingual pretraining corpus of the underlying encoder), with enhanced support for 52 languages and programming code that receive explicit retrieval-pair and cross-lingual training. All training data uses permissive, enterprise-friendly licenses, plus IBM-collect

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite guardian 4.1 8b
Ibm-granite · 8.4B

Granite Guardian 4.1 8B introduces improved Bring Your Own Criteria (BYOC) support, enabling users to define arbitrary judging criteria beyond the pre-baked safety and hallucination detectors. The model can now faithfully evaluate complex, multi-part requirements such as formatting rules, length constraints, and domain-specific instructions.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite vision 4.1 4b
Ibm-granite · 4B

Model Summary: Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint, providing a lightweight alternative to much larger frontier models for these tasks:

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite speech 4.1 2b plus
Ibm-granite · 2.1B

Granite-Speech-4.1-2B-Plus has similar capabilities to the Granite-Speech-4.1-2B model. The plus model adds two new community-requested rich transcription features that can be activated with a simple prompt change: speaker-attributed ASR (speaker labels and word transcripts) and word-level timing information. Unlike the base mode, the plus model doesn't provide punctuation and capitalization.

4K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite speech 4.1 2b
Ibm-granite · 2.3B

Model Summary: Granite Speech 4.1 2B is a compact and efficient speech-language model, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST) for English, French, German, Spanish, Portuguese and Japanese.

4K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B Long Base
LumiOpen · 8B

Poro 2 Long Base is an 8B parameter decoder-only transformer created by extending Poro 2 8B Base from an 8K to 128K token context window using LongRoPE. The model supports both English and Finnish with an extended context window of 128K tokens. Poro 2 Long Base is released as a fully open source model under the Llama 3.1 Community License.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 4B
Swiss-ai · 3.8B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
URSA 1.7B IBQ512
BAAI · 2.3B

Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.1 30b
Ibm-granite · 29B

Model Summary: Granite-4.1-30B is a 30B parameter long-context instruct model finetuned from Granite-4.1-30B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.1 8b base
Ibm-granite · 8.4B

Model Summary: Granite‑4.1‑8B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of text‑to‑text generation tasks. In addition to standard generation, it supports Fill‑in‑the‑Middle (FIM) code completion through specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens using a five‑phase training strategy: 10 trillion tokens in phase one, 2 trillion tokens each in phases two and three, and 0.5 trillion tokens in phase four. In the final phase, long‑context extension is applied to expand the m

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.1 8b
Ibm-granite · 8.8B

Model Summary: Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from Granite-4.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.1 3b base
Ibm-granite · 3.4B

Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture and weights as Granite 4.0 3B Micro, which is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the third, and 0.5 trillion in the final stage. An additional training phase is applied

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.1 3b
Ibm-granite · 3.4B

Model Summary: Granite-4.1-3B is a 3B parameter long-context instruct model finetuned from Granite-4.1-3B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EGM 4B SFT
NVIDIA · 0B

EGM-Qwen3-VL-4B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-4B-Thinking.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EGM 8B SFT
NVIDIA · 0B

EGM-Qwen3-VL-8B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-8B-Thinking.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EGM 4B
NVIDIA · 4.8B

EGM-Qwen3-VL-4B is an efficient visual grounding model from the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-4B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B Long Math Reasoning SFT Preview
LumiOpen · 8B

Poro 2 8B Math Reasoning SFT Preview is a specialized model focused on mathematical reasoning and problem-solving. This preview model was created through supervised fine-tuning of a context-extended base model. This model excels at mathematical reasoning tasks but is not optimized for general conversational use or other domains.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 0.5B
Swiss-ai · 0.4B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Ising Calibration 1 35B A3B
NVIDIA · 0B

This model is a fine-tuned derivative of Qwen3.5-35B-A3B. Follow the Qwen3.5-35B-A3B serving guide for deployment with vLLM, replacing the model path with nvidia/NVIDIA-Ising-Calibration-1-35B-A3B. Suggested inference settings: temperature=0.2, maxtokens=16384.

262K Kontext GPU ab €5,09/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
harrier oss v1 27b
Microsoft · 27B

harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.

131K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
harrier oss v1 270m
Microsoft · 0.3B

harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
UniRG CXR
Microsoft · 8.8B

We introduce UniRG-CXR, a radiology report generation model that obtains SOTA performance on ReXrank. More details can be found in the paper: Scaling medical imaging report generation with multimodal reinforcement learning

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 4 31B
Google · 31B

gemma 4 31B ist ein quelloffenes Sprachmodell von Google mit 31B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 4 26B A4B it
Google · 27B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

262K Kontext Router €0.29 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
gemma 4 31B it
Google · 31B

Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Super 120B A12B NVFP4
NVIDIA · 67B

Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.

262K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Super 120B A12B BF16
NVIDIA · 124B

Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.

262K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Nano 4B BF16
NVIDIA · 4B

The pretraining data has a cutoff date of September 2024\.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Nemotron 3 Content Safety
NVIDIA · 4.3B

Model Dates: Trained between Oct 2025 and March 2026

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 3b vision
Ibm-granite · 4B

Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EGM 8B
NVIDIA · 8.8B

EGM-Qwen3-VL-8B is the flagship model of the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-8B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 0.8B Base
Qwen · 0.9B

[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 0.8B
Qwen · 0.8B

Qwen3.5 0.8B ist ein quelloffenes Sprachmodell von Qwen mit 0.8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 2B
Qwen · 2B

Qwen3.5 2B ist ein quelloffenes Sprachmodell von Qwen mit 2B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 1b speech
Ibm-granite · 2.3B

Model Summary: Granite-4.0-1b-speech is a compact and efficient speech-language model, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST).

4K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 4B
Qwen · 4B

Qwen3.5 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 9B Base
Qwen · 9.7B

[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3.5 27B
Qwen · 28B

[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3.5 35B A3B
Qwen · 35B

Qwen3.5 35B A3B ist ein quelloffenes Sprachmodell von Qwen mit 35B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
llama nv embed reasoning 3b
NVIDIA · 3.2B

llama-nv-embed-reasoning-3b is a 3.2B-parameter embedding model designed to produce high‑quality sentence and document representations for retrieval, semantic search, and similarity tasks, with a strong focus on reasoning‑heavy content.

131K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Apertus v1.1 1.5B
Swiss-ai · 1.5B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
X Reasoner 7B
Microsoft · 8.3B

We introduce X-Reasoner, a vision-language model posttrained solely on general-domain text for generalizable reasoning, using a twostage approach: an initial supervised fine-tuning phase with distilled long chainof-thoughts, followed by reinforcement learning with verifiable rewards. Experiments show that X-Reasoner successfully transfers reasoning capabilities to both multimodal and out-of-domain settings, outperforming existing state-of-theart models trained with in-domain and multimodal data across various general and medical benchmarks. More details can be found in the paper: X-Reasoner: T

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM OCR
Z.AI · 1.3B

GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, GLM-OCR deliver

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite vision 3.3 2b chart2csv preview
Ibm-granite · 3B

Chart2CSV is a specialized vision-language model fine-tuned for the accurate extraction of tabular data from charts and visualizations. Built on top of ibm-granite/granite-vision-3.3-2b, it produces machine-readable CSV outputs with improved numeric fidelity compared to general-purpose VLMs. The model is trained using code-guided synthetic chart data following the ChartGen methodology, which strengthens factual grounding and reduces hallucination in the Chart-to-CSV task.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 ForcedAligner 0.6B
Qwen · 0.9B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 ASR 1.7B
Qwen · 2.3B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 ASR 0.6B
Qwen · 0.9B

The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
ALIA 40b instruct 2601
BSC-LT · 40B

[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.

164K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
DeepSeek OCR 2
DeepSeek · 3.4B

Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:

8K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EuroLLM 9B Instruct 2512
Utter-project · 9.2B

This is the model card for EuroLLM-9B-Instruct-2512, an improved version of utter-project/EuroLLM-9B-Instruct. In comparison with the previous version, this version includes the long-context extension phase and the revamped post-training recipe from utter-project/EuroLLM-22B-Instruct.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Phi 4 reasoning vision 15B
Microsoft · 15B

Developer: Microsoft Corporation Authorized Representative: Microsoft Ireland Operations Limited, 70 Sir John Rogerson's Quay, Dublin 2, D02 R296, Ireland Release Date: March 4, 2026 License: MIT Parameters: 15B Context Length: 16,384 tokens Inputs: Text and Images Outputs: Text Training GPUs: 240 B200s Training Time: 4 days Training Dates: February 3, 2025 – February 21, 2026 Model Dependencies: Phi-4-Reasoning

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.7 Flash
Z.AI · 31B

GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.

203K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
translategemma 27b it
Google · 29B

translategemma 27b it ist ein multimodales Sprachmodell von Google mit 29B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
translategemma 12b it
Google · 13B

translategemma 12b it ist ein multimodales Sprachmodell von Google mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
translategemma 4b it
Google · 5B

translategemma 4b it ist ein multimodales Sprachmodell von Google mit 5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM Image
Z.AI · 6.9B

GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑imag

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paza whisper large v3 turbo
Microsoft · 0.8B

This model is a fine-tuned version of the openai/whisper-large-v3-turbo model finetuned for automatic speech recognition (ASR) in several Kenyan languages, including Swahili, Kalenjin, Kikuyu, Luo, Maasai and Somali. Whisper is a transformer-based encoder-decoder model that converts raw audio into text. The encoder processes audio inputs as log-Mel spectrograms, capturing acoustic and linguistic features, while the decoder generates text tokens in an autoregressive manner. This design allows the model to handle diverse languages, accents, and noise conditions with strong generalization.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
medgemma 1.5 4b it
Google · 4.3B

medgemma 1.5 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paza Phi 4 multimodal instruct
Microsoft · 5.6B

Fine-tuning was performed on the entire unified multilingual ASR dataset, comprising the mentioned six languages, to encourage cross-lingual generalization. During fine-tuning, only the audio-specific components: audio embedding module, audio encoder, and audio projection layers, were unfrozen and set as trainable, while the rest of the model parameters remained frozen to preserve pretrained language capabilities. Dropout was applied to both the audio encoder and projection layers to regularize training. The model leverages a multimodal processor that handles text tokenization and audio featur

131K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL Embedding 2B
Qwen · 2.1B

The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.

262K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL Embedding 8B
Qwen · 8.1B

The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.

262K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen Image 2512
Qwen · 20B

We are excited to introduce Qwen-Image-2512, the December update of Qwen-Image’s text-to-image foundational model. You are welcome to try the latest model at Qwen Chat. Compared to the base Qwen-Image model released in August, Qwen-Image-2512 features the following key improvements:

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.7
Z.AI

GLM-4.7, your new coding partner, is coming with the following features:

203K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Nano 30B A3B NVFP4
NVIDIA · 18B

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper large v3 LoS punctuated
BSC-LT

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · auf Anfrage
EuroMoE 2.6B A0.6B Instruct 2512
Utter-project · 2.6B

This is the model card for EuroMoE-2.6B-A0.6B-Instruct-2512. You can also check the pre-trained version: EuroMoE-2.6B-A0.6B-2512.

33K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Cosmos Reason2 2B
NVIDIA · 2.4B

Cosmos Reason2 2B ist ein multimodales Sprachmodell von NVIDIA mit 2.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos Reason2 8B
NVIDIA · 8.8B

Cosmos Reason2 8B ist ein multimodales Sprachmodell von NVIDIA mit 8.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper large v3 LoS
BSC-LT

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · auf Anfrage
GLM ASR Nano 2512
Z.AI · 2.3B

GLM-ASR-Nano-2512 is a robust, open-source speech recognition model with 1.5B parameters. Designed for real-world complexity, it outperforms OpenAI Whisper V3 on multiple benchmarks while maintaining a compact size.

8K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AutoGLM Phone 9B Multilingual
Z.AI · 0B

⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AutoGLM Phone 9B
Z.AI · 0B

⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.6V
Z.AI · 108B

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

131K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.6V Flash
Z.AI · 10B

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

131K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron 3 Nano 30B A3B BF16
NVIDIA · 32B

The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
llama nemotron embed vl 1b v2
NVIDIA · 1.7B

llama-nemotron-embed-vl-1b-v2 was developed by NVIDIA for multimodal question-answering retrieval. The model can embed document pages in the form of image, text, or combined image–text inputs. Documents can be retrieved given a user query in text form. The model supports page images containing text, tables, charts, and infographics. We report the evaluation of this model on two internal multimodal retrieval benchmarks, and on the popular ViDoRe V1 and V2 benchmarks and the new Vidore V3 benchmark.

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EuroLLM 22B 2512
Utter-project · 23B

This is the model card for EuroLLM-22B. You can also check the post-trained version: EuroLLM-22B-Instruct-2515.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Ministral 3 3B Instruct 2512 ONNX
Mistral · 3B

[!Tip] This model was contributed by Xenova from Hugging Face. We sincerely appreciate the integration and community collaboration. While preliminary functionality checks have been performed, comprehensive testing has not yet been completed. We recommend you to proceed with caution and conducting your own evaluations for specific use cases. If any issues arise, open a PR/Issue here and we will try to address them promptly.

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Salamandra VL 7B 2512
BSC-LT · 8.9B

Salamandra-VL-7B-2512 is the latest version of the Salamandra vision model family. This version brings significant improvements in both architecture and training data.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Voxtral 4B TTS 2603
Mistral · 4B

Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.

GPU ab €1,68/Std. Sprachausgabe 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
WebVIA Agent
Z.AI · 10B

- Repository: https://github.com/zheny2751-dotcom/WebVIA - Paper: https://arxiv.org/pdf/2511.06251

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
UI2Code N
Z.AI · 10B

- Repository: https://github.com/zai-org/UI2CodeN - Paper: https://arxiv.org/abs/2511.08195

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi K2 Thinking
Moonshotai · 1058B

Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage.

262K Kontext In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
URSA 0.6B FSQ320
BAAI · 0.7B

Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
URSA 0.6B IBQ1024
BAAI · 0.9B

Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Fara 7B
Microsoft · 8.3B

Update: We just released Fara1.5 which improves on Fara-7B dramatically and is available in three model sizes 4B, 9B and 27B!

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi Linear 48B A3B Base
Moonshotai · 49B

Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi Linear 48B A3B Instruct
Moonshotai · 49B

Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 7B Surg CholecT50
NVIDIA · 8.3B

This model is for research and development only. <br

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Glyph
Z.AI · 10B

- Repository: https://github.com/thu-coai/Glyph - Paper: https://arxiv.org/abs/2510.17800

131K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
URSA 1.7B IBQ1024
BAAI · 2.3B

Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
URSA 1.7B FSQ320
BAAI · 2B

Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron Nano 12B v2 VL NVFP4 QAD
NVIDIA · 7.7B

NVIDIA-Nemotron-Nano-VL-12B-V2-FP4-QAD is the quantized version of the NVIDIA Nemotron Nano VL V2 model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Nemotron Nano VL FP4 QAD model is quantized with TensorRT Model Optimizer.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 2B Instruct
Qwen · 2.1B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 32B Instruct
Qwen · 33B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext GPU ab €5,09/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek OCR
DeepSeek · 3.3B

torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation

8K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
llama nemotron embed 1b v2
NVIDIA · 1.2B

The Llama Nemotron Embedding 1B model is optimized for multilingual and cross-lingual text question-answering retrieval with support for long documents (up to 8192 tokens) and dynamic embedding size (Matryoshka Embeddings). This model was evaluated on 26 languages: English, Arabic, Bengali, Chinese, Czech, Danish, Dutch, Finnish, French, German, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Norwegian, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, and Turkish.

131K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NV Reason CXR 3B
NVIDIA · 3.8B

This model is for research and development only.

128K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 8B Instruct
Qwen · 8.8B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 4B Instruct
Qwen · 4.4B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
functiongemma 270m it
Google · 0.3B

functiongemma 270m it ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 350m base
Ibm-granite · 0.4B

Model Summary: Granite-4.0-350M-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 350m
Ibm-granite · 0.4B

Model Summary: Granite-4.0-350M is a lightweight instruct model finetuned from Granite-4.0-350M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 1b base
Ibm-granite · 1.6B

Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the th

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 1b
Ibm-granite · 1.6B

Model Summary: Granite-4.0-1B is a lightweight instruct model finetuned from Granite-4.0-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
llama embed nemotron 8b
NVIDIA · 7.5B

This model achieves state-of-the-art performance on the multilingual MTEB leaderboard as of October 21, 2025.

131K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 Nemotron Nano VL 8B V1 FP4 QAD
NVIDIA · 5.7B

Llama-3.1-Nemotron-Nano-VL-8B-V1-FP4-QAD is the quantized version of the NVIDIA Llama Nemotron Nano VL model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Llama Nemotron Nano VL FP4 QAD model is quantized with TensorRT Model Optimizer.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
UserLM 8b
Microsoft · 8B

Unlike typical LLMs that are trained to play the role of the "assistant" in conversation, we trained UserLM-8b to simulate the “user” role in conversation (by training it to predict user turns in a large corpus of conversations called WildChat). This model is useful in simulating more realistic conversations, which is in turn useful in the development of more robust assistants.

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 30B A3B Instruct
Qwen · 31B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.6
Z.AI

Compared with GLM-4.5, GLM-4.6 brings several key improvements:

203K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 7b instruct tools 16k
BSC-LT · 7.8B

[!WARNING] WARNING: This is a language model that has undergone instruction tuning for conversational settings that exploit function calling capabilities. It has not been aligned with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications. How to use from datetime import datetime from transformers import AutoTokenizer, AutoModelForCausalLM import transformer

16K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3.2 Exp
DeepSeek

We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3.1 Terminus
DeepSeek

This update maintains the model's original capabilities while addressing issues reported by users, including:

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 235B A22B Instruct
Qwen · 235B

Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.

262K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gpt oss safeguard 20b
OpenAI · 22B

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.

131K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gpt oss safeguard 120b
OpenAI · 120B

gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.

131K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 h tiny
Ibm-granite · 6.9B

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 h small
Ibm-granite · 32B

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 micro
Ibm-granite · 3.4B

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 h micro
Ibm-granite · 3.2B

📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi K2 Instruct 0905
Moonshotai

Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2. It is a state-of-the-art mixture-of-experts (MoE) language model, featuring 32 billion activated parameters and a total of 1 trillion parameters.

262K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Apertus 8B 2509
Swiss-ai · 8.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

66K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus 70B 2509
Swiss-ai · 71B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

66K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Apertus 70B Instruct 2509
Swiss-ai · 71B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

66K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
embeddinggemma 300m qat q8 0 unquantized
Google · 0.3B

embeddinggemma 300m qat q8 0 unquantized ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
embeddinggemma 300m qat q4 0 unquantized
Google · 0.3B

embeddinggemma 300m qat q4 0 unquantized ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3.1
DeepSeek

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
TowerVision 9B
Utter-project · 9.7B

TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3.1 Base
DeepSeek · 3.1B

DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Apertus 8B Instruct 2509
Swiss-ai · 8.1B

1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects

66K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
NVIDIA Nemotron Nano 9B v2
NVIDIA · 8.9B

The pretraining data has a cutoff date of September 2024.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
TowerVision 2B
Utter-project · 3B

TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.

8K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
GLM 4.5V
Z.AI · 108B

This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

66K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 270m
Google · 0.3B

gemma 3 270m ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 4B Instruct 2507
Qwen · 4B

Qwen3 4B Instruct 2507 ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gpt oss 20b
OpenAI · 22B

Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.

131K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen Image
Qwen · 20B

Install the latest version of diffusers pip install git+https://github.com/huggingface/diffusers

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3 3 Nemotron Super 49B v1 5 FP8
NVIDIA · 50B

Llama-3.3-Nemotron-Super-49B-v1.5-FP8 is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and agentic tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 Coder 30B A3B Instruct
Qwen · 31B

Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:

262K Kontext Router €0.23 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
gemma 3 270m it
Google · 0.3B

gemma 3 270m it ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Teuken 7B instruct v0.6
OpenGPT-X · 7.5B

- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Wan2.2 T2V A14B Diffusers
Wan-AI · 14B

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.2 TI2V 5B Diffusers
Wan-AI · 5B

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 0528 NVFP4 v2
NVIDIA

Compared to nvidia/DeepSeek-R1-0528-FP4, this checkpoint additionally quantizes the wo module in attention layers.

164K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.5 Base
Z.AI · 4.5B

The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.

131K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.5 Air
Z.AI · 110B

The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.

131K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.5
Z.AI

The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.

131K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.5 Air Base
Z.AI · 110B

The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.

131K Kontext Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.2 TI2V 5B
Wan-AI · 5B

We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding small english r2
Ibm-granite · 0B

Model Summary: Granite-embedding-small-english-r2 is a 47M parameter dense biencoder embedding model from the Granite Embeddings collection that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 based on context length of upto 8192 tokens. Compared to most other open-source models, this model was only trained using open-source relevance-pair datasets with permissive, enterprise-friendly license, plus IBM collected and generated datasets.

8K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding english r2
Ibm-granite · 0.1B

Model Summary: Granite-embedding-english-r2 is a 149M parameter dense biencoder embedding model from the Granite Embeddings collection that can be used to generate high quality text embeddings. This model produces embedding vectors of size 768 based on context length of upto 8192 tokens. Compared to most other open-source models, this model was only trained using open-source relevance-pair datasets with permissive, enterprise-friendly license, plus IBM collected and generated datasets.

8K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
embeddinggemma 300m
Google · 0.3B

embeddinggemma 300m ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper timestamped cs
BSC-LT · 1.5B

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

GPU ab €1,68/Std. Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
whisper 3cat balearic
BSC-LT

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · auf Anfrage
whisper 3cat cv21 valencian
BSC-LT

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · auf Anfrage
MediPhi Instruct
Microsoft · 3.8B

The MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e. PubMed commercial, Medical Wikipedia, Medical Guidelines, Medical Coding, and open-source clinical documents) and merged back with the SLERP method in their base model to conserve general abilities. One model combined all five experts into one general expert with the multi-model merging method BreadCrumbs. F

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi K2 Instruct
Moonshotai

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.

131K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
medgemma 27b it
Google · 29B

medgemma 27b it ist ein multimodales Sprachmodell von Google mit 29B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite docling 258M mlx
Ibm-granite · 0.3B

This model was converted to MLX format from ibm-granite/granite-docling-258M using mlx-vlm version 0.3.3. Refer to the original model card for more details on the model.

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
FLUX.1 Krea dev
Black-forest-labs · 12B

FLUX.1 Krea dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Dayhoff 3b GR HM c
Microsoft · 3B

Dayhoff is an Atlas of both protein sequence data and generative language models — a centralized resource that brings together 3.34 billion protein sequences across 1.7 billion clusters of metagenomic and natural protein sequences (GigaRef), 46 million structure-derived synthetic sequences (BackboneRef), and 16 million multiple sequence alignments (OpenProteinSet). These models can natively predict zero-shot mutation effects on fitness, scaffold structural motifs by conditioning on evolutionary or structural context, and perform guided generation of novel proteins within specified families. Le

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi K2 Base
Moonshotai · 2B

Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.

131K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.1V 9B Thinking
Z.AI · 10B

Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4.1V 9B Base
Z.AI · 10B

Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.

66K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos Predict2 0.6B Text2Image
NVIDIA · 0.6B

Cosmos Predict2 0.6B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 0.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi tiny MoE instruct
Microsoft · 3.8B

Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a larger variant, Phi-mini-MoE, with 7.6B total and 2.4B activated pa

4K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi mini MoE instruct
Microsoft · 7.6B

Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a smaller variant, Phi-tiny-MoE, with 3.8B total and 1.1B activated p

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi VL A3B Thinking 2506
Moonshotai · 16B

[!Note] This is an improved version of Kimi-VL-A3B-Thinking. Please consider using this updated model instead of the previous version.

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 mini flash reasoning
Microsoft · 3.9B

Phi-4-mini-flash-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 64K token context length.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi Dev 72B
Moonshotai · 73B

We introduce Kimi-Dev-72B, our new open-source coding LLM for software engineering tasks. Kimi-Dev-72B achieves a new state-of-the-art on SWE-bench Verified among open-source models.

131K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 70B SFT
LumiOpen · 71B

Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 70B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.

8K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B SFT
LumiOpen · 8B

Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 8B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 3n E2B it
Google · 5.4B

gemma 3n E2B it ist ein multimodales Sprachmodell von Google mit 5.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EuroMoE 2.6B A0.6B Instruct Preview
Utter-project · 2.6B

⚠️ PREVIEW RELEASE: This is a preview version of EuroMoE-2.6B-A0.6B-Instruct-Preview. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroMoE 2.6B A0.6B 2512
Utter-project · 2.6B

This is the model card for EuroLLM-2.6B-A0.6-2512, the pre-trained model for EuroLLM-2.6B-A0.6-2512-Instruct.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroLLM 22B Instruct Preview
Utter-project · 23B

This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.

4K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroVLM 1.7B Preview
Utter-project · 2.1B

⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-1.7B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroVLM 9B Preview
Utter-project · 9.6B

⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-9B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroLLM 9B 2512
Utter-project · 9.2B

This is the model card for EuroLLM-9B-2512, an improved version of utter-project/EuroLLM-9B. In comparison with the previous version, this version includes the long-context extension phase from utter-project/EuroLLM-22B.

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 3n E4B it
Google · 7.8B

gemma 3n E4B it ist ein multimodales Sprachmodell von Google mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 Embedding 4B
Qwen · 4B

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code re

41K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 Embedding 0.6B
Qwen · 0.6B

The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code re

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GUI Actor Verifier 2B
Microsoft · 2.2B

This model was introduced in the paper GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. It is developed based on UI-TARS-2B-SFT and is designed to predict the correctness of an action position given a language instruction. This model is well-suited for GUI-Actor, as its attention map effectively provides diverse candidates for verification with only a single inference.

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 70B base
LumiOpen · 71B

Poro 2 70B Base is a 70B parameter decoder-only transformer created through continued pretraining of Llama 3.1 70B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.

8K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B base
LumiOpen · 8B

Poro 2 8B Base is an 8B parameter decoder-only transformer created through continued pretraining of Llama 3.1 8B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 70B Instruct
LumiOpen · 71B

Poro 2 70B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 70B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.

8K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Poro 2 8B Instruct
LumiOpen · 8B

Poro 2 8B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 8B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 7b vision
BSC-LT · 8.2B

[!WARNING] WARNING: This model has been deprecated and is no longer recommended. For the latest model, please visit: https://huggingface.co/BSC-LT/Salamandra-VL-7B-2512

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
medgemma 27b text it
Google · 27B

medgemma 27b text it ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
medgemma 4b it
Google · 4.3B

medgemma 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite docling 258M
Ibm-granite · 0.3B

Granite Docling 258M builds upon the Idefics3 architecture, but introduces two key modifications: it replaces the vision encoder with siglip2-base-patch16-512 and substitutes the language model with a Granite 165M LLM. Try out our Granite-Docling-258 demo today.

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL v1.5 mmeb
BAAI · 7.6B

2025-4-2 🌟🌟 BGE-VL models are also available on WiseModel.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL Screenshot
BAAI · 3.8B

2025-04-06 🚀🚀 MVRB Dataset are released on Huggingface: MVRB

128K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge code v1
BAAI · 1.5B

For more details please refer to our Github: FlagEmbedding.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL v1.5 zs
BAAI · 7.6B

2025-4-2 🌟🌟 BGE-VL models are also available on WiseModel.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper large v3 ca punctuated 3370h
BSC-LT

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · auf Anfrage
whisper bsc large v3 cat
BSC-LT · 1.5B

- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information

GPU ab €1,68/Std. Transkription 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
NextCoder 7B
Microsoft · 7.6B

NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits (ICML'2025)

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 4.0 tiny preview
Ibm-granite · 6.7B

Model Summary: Granite-4-Tiny-Preview is a 7B parameter fine-grained hybrid mixture-of-experts (MoE) instruct model fine-tuned from Granite-4.0-Tiny-Base-Preview using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised fine-tuning, and model alignment using reinforcement learning.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 mini reasoning
Microsoft · 3.8B

Phi-4-mini-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 128K token context length.

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite speech 3.3 2b
Ibm-granite · 3B

Model Summary: Granite-speech-3.3-2b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.3-2b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.3-2b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.

131K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 235B A22B
Qwen · 235B

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:

41K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 32B
Qwen · 32B

Qwen3 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 30B A3B
Qwen · 30B

Qwen3 30B A3B ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 14B
Qwen · 14B

Qwen3 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 8B
Qwen · 8B

Qwen3 8B ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 4B
Qwen · 4B

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:

41K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 1.7B
Qwen · 1.7B

Qwen3 1.7B ist ein quelloffenes Sprachmodell von Qwen mit 1.7B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 0.6B
Qwen · 0.8B

Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:

41K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Kimi Audio 7B
Moonshotai · 9.8B

We present Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation. This repository hosts the model checkpoints for Kimi-Audio-7B.

8K Kontext GPU ab €1,68/Std. Sprachausgabe In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi Audio 7B Instruct
Moonshotai · 9.8B

We present Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation. This repository hosts the model checkpoints for Kimi-Audio-7B-Instruct.

8K Kontext GPU ab €1,68/Std. Sprachausgabe In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama Guard 4 12B
Meta · 12B

Llama Guard 4 12B ist ein multimodales Sprachmodell von Meta mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos Predict2 14B Text2Image
NVIDIA · 14B

Cosmos Predict2 14B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 14B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos Predict2 2B Text2Image
NVIDIA · 2B

Cosmos Predict2 2B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos Reason1 7B
NVIDIA · 8.3B

NVIDIA Cosmos Reason – an open, customizable, 7B-parameter reasoning vision language model (VLM) for physical AI and robotics - enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 reasoning plus
Microsoft · 15B

[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite speech 3.3 8b
Ibm-granite · 8.6B

Model Summary: Granite-speech-3.3-8b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.3-8b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.3-8b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.

131K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM Z1 Rumination 32B 0414
Z.AI · 33B

The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 reasoning
Microsoft · 15B

[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.3 8b instruct
Ibm-granite · 8.2B

Model Summary: Granite-3.3-8B-Instruct is a 8-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-8B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.3 2b instruct
Ibm-granite · 2.5B

Model Summary: Granite-3.3-2B-Instruct is a 2-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-2B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi VL A3B Thinking
Moonshotai · 16B

[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kimi VL A3B Instruct
Moonshotai · 16B

We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities—all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B).

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 12b it qat q4 0 unquantized
Google · 12B

gemma 3 12b it qat q4 0 unquantized ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM Z1 32B 0414
Z.AI · 33B

The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM Z1 9B 0414
Z.AI · 9.4B

The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4 32B Base 0414
Z.AI · 33B

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4 32B 0414
Z.AI · 33B

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
GLM 4 9B 0414
Z.AI · 9.4B

The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 4 Maverick 17B 128E
Meta · 17B

Llama 4 Maverick 17B 128E ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 4 Scout 17B 16E
Meta · 109B

Llama 4 Scout 17B 16E ist ein multimodales Sprachmodell von Meta mit 109B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €1.20 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 4 Scout 17B 16E Instruct
Meta · 109B

Llama 4 Scout 17B 16E Instruct ist ein multimodales Sprachmodell von Meta mit 109B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €1.20 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 4 Maverick 17B 128E Instruct
Meta · 17B

Llama 4 Maverick 17B 128E Instruct ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite speech 3.2 8b
Ibm-granite · 8.5B

Model Summary: Granite-speech-3.2-8b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.2-8b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.2-8b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.

131K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3 0324
DeepSeek

DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
txgemma 2b predict
Google · 2.6B

txgemma 2b predict ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 32B Instruct
Qwen · 33B

In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.

128K Kontext GPU ab €5,09/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 7b instruct tools
BSC-LT · 7.8B

salamandra 7b instruct tools ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 Nemotron Nano 8B v1
NVIDIA · 8B

Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3 3 Nemotron Super 49B v1
NVIDIA · 50B

Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.

131K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 1b it
Google · 1B

gemma 3 1b it ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL MLLM S2
BAAI · 7.6B

2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL MLLM S1
BAAI · 7.6B

2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CogView4 6B
Z.AI · 6.4B

+ Resolution: Width and height must be between 512px and 2048px, divisible by 32, and ensure the maximum number of pixels does not exceed 2^21 px. + Precision: BF16 / FP32 (FP16 is not supported as it will cause overflow resulting in completely black images)

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 12b it
Google · 12B

gemma 3 12b it ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 12b pt
Google · 12B

gemma 3 12b pt ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 27b it
Google · 27B

gemma 3 27b it ist ein multimodales Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.10 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
gemma 3 27b pt
Google · 27B

gemma 3 27b pt ist ein multimodales Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.1 T2V 14B Diffusers
Wan-AI · 14B

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.1 T2V 1.3B Diffusers
Wan-AI · 1.4B

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.1 T2V 1.3B
Wan-AI · 1.4B

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan2.1 T2V 14B
Wan-AI · 14B

In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL large
BAAI · 0.4B

2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.

77 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BGE VL base
BAAI · 0.1B

2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.

77 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 multimodal instruct
Microsoft · 5.6B

🎉Phi-4: [mini-reasoning | reasoning] | [multimodal-instruct | onnx]; [mini-instruct | onnx]

131K Kontext GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Moonlight 16B A3B Instruct
Moonshotai · 16B

- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Moonlight 16B A3B
Moonshotai · 16B

- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 1b pt
Google · 1B

gemma 3 1b pt ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 4b it
Google · 4.3B

gemma 3 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 4b pt
Google · 4.3B

gemma 3 4b pt ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 mini instruct
Microsoft · 3.8B

🎉Phi-4: [mini-reasoning | reasoning] | [multimodal-instruct | onnx]; [mini-instruct | onnx]

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 30m sparse
Ibm-granite · 0B

Model Summary: Granite-Embedding-30m-Sparse is a 30M parameter sparse biencoder embedding model from the Granite Experimental suite that can be used to generate high quality text embeddings. This model produces variable length bag-of-word like dictionary, containing expansions of sentence tokens and their corresponding weights and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. While maintaining competitive scores on academic benchmarks such as BEIR, this model also performs well on ma

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.2 2b instruct
Ibm-granite · 2.5B

Model Summary: Granite-3.2-2B-Instruct is an 2-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-2B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic data designed for reasoning tasks. The model allows controllability of its thinking capability, ensuring it is applied only when required.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite vision 3.2 2b
Ibm-granite · 3B

Model Summary: granite-vision-3.2-2b is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model with both image and text modalities.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite vision 3.1 2b preview
Ibm-granite · 3B

Model Summary: granite-vision-3.1-2b-preview is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model (https://huggingface.co/ibm-granite/granite-3.1-2b-instruct) with

16K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 72B Instruct
Qwen · 73B

--- license: other licensename: qwen licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-72B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---

128K Kontext Router €0.26 pro 1M Input GPU ab €5,09/Std. Vision In der EU gehostet Router · warm Dedicated · verfügbar
Qwen2.5 VL 7B Instruct
Qwen · 8.3B

--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---

128K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 3B Instruct
Qwen · 3.8B

--- licensename: qwen-research licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---

128K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Qwen 32B
DeepSeek · 33B

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Qwen 14B
DeepSeek · 15B

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Qwen 7B
DeepSeek · 7.6B

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Llama 70B
DeepSeek · 70B

DeepSeek R1 Distill Llama 70B ist ein quelloffenes Sprachmodell von DeepSeek mit 70B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Llama 8B
DeepSeek · 8B

DeepSeek R1 Distill Llama 8B ist ein quelloffenes Sprachmodell von DeepSeek mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill Qwen 1.5B
DeepSeek · 1.8B

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1
DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Zero
DeepSeek

We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm 4 9b hf
Z.AI · 9.4B

If you are using the weights from this repository, please update to

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Cosmos 1.0 Diffusion 7B Text2World
NVIDIA · 7B

Cosmos 1.0 Diffusion 7B Text2World ist ein quelloffenes Sprachmodell von NVIDIA mit 7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
OmniGen v1
BAAI · 3.9B

More information please refer to our repo: https://github.com/VectorSpaceLab/OmniGen

131K Kontext GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3
DeepSeek

We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
cogagent 9b 20241220
Z.AI · 14B

The CogAgent-9B-20241220 model is based on GLM-4V-9B, a bilingual open-source VLM base model. Through data collection and optimization, multi-stage training, and strategy improvements, CogAgent-9B-20241220 achieves significant advancements in GUI perception, inference prediction accuracy, action space completeness, and task generalizability. The model supports bilingual (Chinese and English) interaction with both screenshots and language input.

Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite guardian 3.1 2b
Ibm-granite · 2.5B

Granite Guardian 3.1 2B is a fine-tuned Granite 3.1 2B Instruct model designed to detect risks in prompts and responses. It can help with risk detection along many key dimensions catalogued in the IBM AI Risk Atlas. It is trained on unique data comprising human annotations and synthetic data informed by internal red-teaming. It outperforms other open-source models in the same space on standard benchmarks.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
nova d48w1024 osp480
BAAI · 0.6B

Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
nova d48w1024 sd512
BAAI · 0.6B

Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
nova d48w1536 sdxl1024
BAAI · 1.5B

Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
nova d48w1024 sdxl1024
BAAI · 0.6B

Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
nova d48w768 sdxl1024
BAAI · 0.4B

Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
phi 4
Microsoft · 15B

Our training data is an extension of the data used for Phi-3 and includes a wide variety of sources from:

16K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Teuken 7B base v0.6
OpenGPT-X · 7.5B

- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
ALIA 40b
BSC-LT · 40B

[!WARNING] WARNING: This is a base language model that has not undergone instruction tuning or alignment with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications.

33K Kontext GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
granite 3.1 1b a400m instruct
Ibm-granite · 1.3B

Model Summary: Granite-3.1-1B-A400M-Instruct is a 1B parameter long-context instruct model finetuned from Granite-3.1-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.1 3b a800m instruct
Ibm-granite · 3.3B

Model Summary: Granite-3.1-3B-A800M-Instruct is a 3B parameter long-context instruct model finetuned from Granite-3.1-3B-A800M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.1 2b instruct
Ibm-granite · 2.5B

Model Summary: Granite-3.1-2B-Instruct is a 2B parameter long-context instruct model finetuned from Granite-3.1-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.1 8b instruct
Ibm-granite · 8.2B

Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Aquila 135M Instruct
BAAI · 0.3B

Aquila 135M Instruct ist ein quelloffenes Sprachmodell von BAAI mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 278m multilingual
Ibm-granite · 0.3B

Model Summary: Granite-Embedding-278M-Multilingual is a 278M parameter model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 768 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. This model is developed using contrastive finetuning, knowledge distillation and model merging for improved performance.

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 107m multilingual
Ibm-granite · 0.1B

Model Summary: Granite-Embedding-107M-Multilingual is a 107M parameter dense biencoder embedding model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. This model is developed using contrastive finetuning, knowledge distillation and model merging for improved performance.

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 30m english
Ibm-granite · 0B

Model Summary: Granite-Embedding-30m-English is a 30M parameter dense bi-encoder embedding model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. While maintaining competitive scores on academic benchmarks such as BEIR, this model also performs well on many enterprise use cases. This model is developed using retrieval oriented pre-training, c

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite embedding 125m english
Ibm-granite · 0.1B

News: Granite Embedding R2 models with 8192 context length released.

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 large controlnet depth
Stabilityai · 2.2B

This repository provides the Depth ControlNet for Stable Diffusion 3.5 Large..

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 large controlnet blur
Stabilityai · 2.2B

This repository provides the Blur ControlNet for Stable Diffusion 3.5 Large..

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 large controlnet canny
Stabilityai · 2.2B

This repository provides the Canny ControlNet for Stable Diffusion 3.5 Large..

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm edge v 5b
Z.AI · 4.9B

Install the transformers library from the source code:

4K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm edge v 2b
Z.AI · 2.1B

Install the transformers library from the source code:

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
EuroLLM 9B Instruct
Utter-project · 9.2B

EuroLLM 9B Instruct ist ein quelloffenes Sprachmodell von Utter-project mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroLLM 9B
Utter-project · 9.2B

EuroLLM 9B ist ein quelloffenes Sprachmodell von Utter-project mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
paligemma2 3b pt 896
Google · 3B

paligemma2 3b pt 896 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma2 3b pt 448
Google · 3B

paligemma2 3b pt 448 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma2 3b pt 224
Google · 3B

paligemma2 3b pt 224 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma2 3b ft docci 448
Google · 3B

paligemma2 3b ft docci 448 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma2 3b mix 224
Google · 3B

paligemma2 3b mix 224 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
FLUX.1 Depth dev
Black-forest-labs · 12B

FLUX.1 Depth dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
FLUX.1 Canny dev
Black-forest-labs · 12B

FLUX.1 Canny dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm edge 4b chat
Z.AI · 4.3B

Install the transformers library from the source code:

8K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm edge 1.5b chat
Z.AI · 1.6B

Install the transformers library from the source code:

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 7b instruct aina hack
BSC-LT · 7.8B

Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 7B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 2b instruct aina hack
BSC-LT · 2.3B

Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 2B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 Coder 32B Instruct
Qwen · 33B

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 Coder 14B Instruct
Qwen · 15B

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CogVideoX1.5 5B
Z.AI · 5.6B

CogVideoX is an open-source video generation model similar to QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 2b base gptq
BSC-LT · 2.3B

This model is the gptq-quantized version of Salamandra-2b for speculative decoding.

8K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 7b base gptq
BSC-LT · 7.8B

This model is the gptq-quantized version of Salamandra-7b for speculative decoding.

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 medium
Stabilityai · 2.5B

stable diffusion 3.5 medium ist ein multimodales Sprachmodell von Stabilityai mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Teuken 7B instruct commercial v0.4
OpenGPT-X · 7.5B

- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Emu3 Gen hf
BAAI · 8.8B

Below is the model card of Emu3-Chat model, which is adapted from the original Emu3 model card that you can find here.

9K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Emu3 Chat hf
BAAI · 8.8B

Below is the model card of Emu3-Chat model, which is adapted from the original Emu3 model card that you can find here.

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm 4 9b chat 1m hf
Z.AI · 9.5B

If you are using the weights from this repository, please update to

1M Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm 4 9b chat hf
Z.AI · 9.4B

If you are using the weights from this repository, please update to

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 large turbo
Stabilityai · 8.1B

stable diffusion 3.5 large turbo ist ein multimodales Sprachmodell von Stabilityai mit 8.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3.5 large
Stabilityai · 8.1B

stable diffusion 3.5 large ist ein multimodales Sprachmodell von Stabilityai mit 8.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
OmniParser
Microsoft

This model hub includes a finetuned version of YOLOv8 and a finetuned BLIP-2 model on the above dataset respectively. For more details of the models used and finetuning, please refer to the paper.

Router €0.05 pro 1M Input Vision In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
CogView3 Plus 3B
Z.AI · 2.8B

This model is the DiT version of CogView3, a text-to-image generation model, supporting image generation from 512 to 2048px.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.0 1b a400m instruct
Ibm-granite · 1.3B

Model Summary: Granite-3.0-1B-A400M-Instruct is an 1B parameter model finetuned from Granite-3.0-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.0 1b a400m base
Ibm-granite · 1.4B

Model Summary: Granite-3.0-1B-A400M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 8 trillion tokens sourced from diverse domains. During the second stage, it is further trained on 2 trillion tokens using a carefully curated mix of high-quality data, aiming to enhance its performance on specific tasks.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.0 8b instruct
Ibm-granite · 8.2B

Model Summary: Granite-3.0-8B-Instruct is a 8B parameter model finetuned from Granite-3.0-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.0 8b base
Ibm-granite · 8.2B

Model Summary: Granite-3.0-8B-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 10 trillion tokens sourced from diverse domains. During the second stage, it is further trained on 2 trillion tokens using a carefully curated mix of high-quality data, aiming to enhance its performance on specific tasks.

4K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3.0 2b instruct
Ibm-granite · 2.6B

Model Summary: Granite-3.0-2B-Instruct is a 2B parameter model finetuned from Granite-3.0-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper large v3 turbo
OpenAI · 0.8B

Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on 5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
NVLM D 72B
NVIDIA · 79B

This model is ready for non-commercial use.

Router €1.20 pro 1M Input GPU ab €9,13/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra 2b instruct
BSC-LT · 2.3B

This repository contains the model described in Salamandra Technical Report.

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 2b
BSC-LT · 2.3B

This repository contains the model described in Salamandra Technical Report.

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 7b instruct
BSC-LT · 7.8B

This repository contains the model described in Salamandra Technical Report.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
salamandra 7b
BSC-LT · 7.8B

This repository contains the model described in Salamandra Technical Report.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 2 2b jpn it
Google · 2.6B

gemma 2 2b jpn it ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Checkpoint 4epoch rag
BSC-LT · 7.8B

--- license: apache-2.0 datasets: - projecte-aina/RAGMultilingual language: - es - en - ca libraryname: transformers ---

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Teuken 7B instruct research v0.4
OpenGPT-X · 7.5B

- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Flor 6.3B Instruct 4096
BSC-LT · 6.2B

--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---

Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Guard 3 1B
Meta · 1.5B

Llama Guard 3 1B ist ein quelloffenes Sprachmodell von Meta mit 1.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 3B
Meta · 3B

Llama 3.2 3B ist ein quelloffenes Sprachmodell von Meta mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 3B Instruct
Meta · 3.2B

Llama 3.2 3B Instruct ist ein quelloffenes Sprachmodell von Meta mit 3.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 1B Instruct
Meta · 1.2B

Llama 3.2 1B Instruct ist ein quelloffenes Sprachmodell von Meta mit 1.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Flor 6.3B Instruct
BSC-LT · 6.2B

--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---

Router €0.03 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 1.5B Instruct
Qwen · 1.5B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 3B Instruct
Qwen · 3.1B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 Coder 7B Instruct
Qwen · 7.6B

Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 32B Instruct
Qwen · 33B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 14B Instruct
Qwen · 15B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 7B Instruct
Qwen · 7.6B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 0.5B Instruct
Qwen · 0.5B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 1.5B
Qwen · 1.5B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 0.5B
Qwen · 0.5B

Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
OPI Llama 3.1 8B Instruct
BAAI · 8B

OPI Llama 3.1 8B Instruct ist ein quelloffenes Sprachmodell von BAAI mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Checkpoint 2b instructed beta
BSC-LT · 2.3B

2b Version of model Sxxxxxx, without last epoch, instructed with baseline dataset including RAGMultilingual

8K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Qwen2 VL 7B Instruct
Qwen · 8.3B

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2 VL 2B Instruct
Qwen · 2.2B

We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3.5 MoE instruct
Microsoft · 42B

Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures.

131K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CogVideoX 5b
Z.AI · 5.6B

CogVideoX is an open-source version of the video generation model originating from QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3.5 vision instruct
Microsoft · 4.1B

Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 model family, and the multimodal version comes with 128K context length (in tokens) it can support. The model underwent a rigorous enhancement process, incorporating both supervised fine-tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures.

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3.5 mini instruct
Microsoft · 3.8B

🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LongWriter glm4 9b
Z.AI · 9.4B

LongWriter-glm4-9b is trained based on glm-4-9b, and is capable of generating 10,000+ words at once.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
salamandra7b rag prompt ca en es
BSC-LT · 7.8B

This instructed model uses a chat template that must be adhered to the input for conversational use. The easiest way to apply it is using the tokenizer's built-in chat template, as shown in the following snippet.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroLLM 1.7B Instruct
Utter-project · 1.7B

This is the model card for the first instruction tuned model of the EuroLLM series: EuroLLM-1.7B-Instruct. You can also check the pre-trained version: EuroLLM-1.7B.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
EuroLLM 1.7B
Utter-project · 1.7B

This is the model card for the first pre-trained model of the EuroLLM series: EuroLLM-1.7B. You can also check the instruction tuned version: EuroLLM-1.7B-Instruct.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
CogVideoX 2b
Z.AI · 1.7B

CogVideoX is an open-source version of the video generation model originating from QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
FLUX.1 dev
Black-forest-labs · 12B

FLUX.1 dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
ar stablelm 2 base
Stabilityai · 1.6B

ar stablelm 2 base ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge en icl
BAAI · 7.1B

For more details please refer to our Github: FlagEmbedding.

33K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 7M Gen Llama3 1 70B
BAAI · 71B

Infinity-Instruct-7M-Gen-Llama3.1-70B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-7M and Infinity-Instruct-Gen and showing favorable results on AlpacaEval 2.0 and arena-hard compared to GPT4.

8K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
experimental7b rag instruct
BSC-LT · 7.8B

experimental7b rag instruct ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern und einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama Guard 3 8B
Meta · 8B

Llama Guard 3 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
experimental7b rag
BSC-LT · 7.8B

experimental7b rag ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern und einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 8B Instruct
Meta · 8B

Llama 3.1 8B Instruct ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
shieldgemma 9b
Google · 9.2B

shieldgemma 9b ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
shieldgemma 2b
Google · 2.6B

shieldgemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 405B Instruct
Meta · 405B

Llama 3.1 405B Instruct ist ein quelloffenes Sprachmodell von Meta mit 405B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 70B Instruct
Meta · 71B

Llama 3.1 70B Instruct ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 2b it
Google · 2.6B

gemma 2 2b it ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 2b
Google · 2.6B

gemma 2 2b ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 405B
Meta · 405B

Llama 3.1 405B ist ein quelloffenes Sprachmodell von Meta mit 405B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 70B
Meta · 70B

Llama 3.1 70B ist ein quelloffenes Sprachmodell von Meta mit 70B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.1 8B
Meta · 8B

Llama 3.1 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0625 Llama3 8B
BAAI · 8B

Infinity-Instruct-3M-0625-Llama3-8B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0625 Yi 1.5 9B
BAAI · 8.8B

Infinity-Instruct-3M-0625-Yi-1.5-9B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0625 Qwen2 7B
BAAI · 7.6B

Infinity-Instruct-3M-0625-Qwen2-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.

131K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0625 Mistral 7B
BAAI · 7.2B

Infinity-Instruct-3M-0625-Mistral-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 compared to Mixtral 8x7B v0.1, Gemini Pro, and GPT-3.5.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codegeex4 all 9b
Z.AI · 9.4B

We introduce CodeGeeX4-ALL-9B, the open-source version of the latest CodeGeeX4 model series. It is a multilingual code generation model continually trained on the GLM-4-9B, significantly enhancing its code generation capabilities. Using a single CodeGeeX4-ALL-9B model, it can support comprehensive functions such as code completion and generation, code interpreter, web search, function call, repository-level code Q&A, covering various scenarios of software development. CodeGeeX4-ALL-9B has achieved highly competitive performance on public benchmarks, such as BigCodeBench and NaturalCodeBench. I

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0613 Llama3 70B
BAAI · 71B

Infinity-Instruct-3M-0613-Llama3-70B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0613 and showing favorable results on AlpacaEval 2.0 compared to GPT4-0613.

8K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 9b
Google · 9.2B

gemma 2 9b ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 9b it
Google · 9.2B

gemma 2 9b it ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 27b
Google · 27B

gemma 2 27b ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2 27b it
Google · 27B

gemma 2 27b it ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Infinity Instruct 3M 0613 Mistral 7B
BAAI · 7.2B

Infinity-Instruct-3M-0613-Mistral-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0613 and showing favorable results on AlpacaEval 2.0 compared to Mixtral 8x7B v0.1, Gemini Pro, and GPT-3.5.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek Coder V2 Lite Instruct
DeepSeek · 16B

In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.

164K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek Coder V2 Instruct
DeepSeek

In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.

164K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion 3 medium diffusers
Stabilityai · 2.1B

stable diffusion 3 medium diffusers ist ein multimodales Sprachmodell von Stabilityai mit 2.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
glm 4 9b
Z.AI · 9.4B

2024/08/12, 本仓库代码已更新并使用 transformers=4.44.0, 请及时更新依赖。

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
phi 2 pytdml
Microsoft · 2.8B

Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-2 showcased a nearly state-of-the-art performance among models with less than 13 billion parameters.

2K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2 1.5B Instruct
Qwen · 1.5B

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the instruction-tuned 1.5B Qwen2 model.

33K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2 0.5B
Qwen · 0.5B

Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language model.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 vision 128k instruct
Microsoft · 4.1B

🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V2 Lite Chat
DeepSeek · 16B

Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:

164K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V2 Lite
DeepSeek · 16B

Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:

164K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LLARA pretrain
BAAI · 6.7B

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

4K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LLARA beir
BAAI · 6.7B

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

4K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LLARA document
BAAI · 6.7B

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

4K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LLARA passage
BAAI · 6.7B

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

4K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma 3b ft cococap 448
Google · 2.9B

paligemma 3b ft cococap 448 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma 3b mix 224
Google · 2.9B

paligemma 3b mix 224 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
paligemma 3b pt 224
Google · 2.9B

paligemma 3b pt 224 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. Vision In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 small 8k instruct
Microsoft · 7.4B

🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

8K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 medium 128k instruct
Microsoft · 14B

🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

131K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 medium 4k instruct
Microsoft · 14B

🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

4K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm 2 instruct 1 6b
Stabilityai · 1.6B

japanese stablelm 2 instruct 1 6b ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm 2 base 1 6b
Stabilityai · 1.6B

japanese stablelm 2 base 1 6b ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
dragon multiturn context encoder
NVIDIA

tokenizer = AutoTokenizer.frompretrained('nvidia/dragon-multiturn-query-encoder') queryencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-query-encoder') contextencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-context-encoder')

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
dragon multiturn query encoder
NVIDIA

tokenizer = AutoTokenizer.frompretrained('nvidia/dragon-multiturn-query-encoder') queryencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-query-encoder') contextencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-context-encoder')

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
DeepSeek V2 Chat
DeepSeek

Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.

164K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 20b code instruct 8k
Ibm-granite · 20B

New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.

Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 8b code instruct 4k
Ibm-granite · 8.1B

New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3b code instruct 2k
Ibm-granite · 3.5B

New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.

2K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
granite 3b code base 2k
Ibm-granite · 3.5B

New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.

2K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 mini 128k instruct
Microsoft · 3.8B

🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 3 mini 4k instruct
Microsoft · 3.8B

🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)

4K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V2
DeepSeek

Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.

164K Kontext Router €0.40 pro 1M Input GPU ab €21,48/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Meta Llama Guard 2 8B
Meta · 8B

Meta Llama Guard 2 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Meta Llama 3 8B
Meta · 8B

Meta Llama 3 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Meta Llama 3 8B Instruct
Meta · 8B

Meta Llama 3 8B Instruct ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Meta Llama 3 70B Instruct
Meta · 71B

Meta Llama 3 70B Instruct ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Meta Llama 3 70B
Meta · 71B

Meta Llama 3 70B ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 2 1 6b chat
Stabilityai · 1.6B

Stable LM 2 Chat 1.6B is a 1.6 billion parameter instruction tuned language model inspired by HugginFaceH4's Zephyr 7B training pipeline. The model is trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 2 12b chat
Stabilityai · 12B

Stable LM 2 12B Chat is a 12 billion parameter instruction tuned language model trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).

4K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Poro 34B chat
LumiOpen · 34B

Poro 34b chat is a chat-tuned version of Poro 34B trained to follow instructions in both Finnish and English. Quantized versions are available on Poro 34B-chat-GGUF.

GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 1.1 2b it
Google · 2.5B

gemma 1.1 2b it ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 1.1 7b it
Google · 8.5B

gemma 1.1 7b it ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 2 12b
Stabilityai · 12B

Stable LM 2 12B is a 12.1 billion parameter decoder-only language model pre-trained on 2 trillion tokens of diverse multilingual and code datasets for two epochs.

4K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codegemma 2b
Google · 2.5B

codegemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
tiny random stablelm 2
Stabilityai · 0.1B

This repository stores a development version of Stable LM 2 for sanity-checking/debugging the transformers implementation.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 34b Instruct hf
Meta · 34B

CodeLlama 34b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 34B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 70b Instruct hf
Meta · 69B

CodeLlama 70b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 13b Instruct hf
Meta · 13B

CodeLlama 13b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 13b Python hf
Meta · 13B

CodeLlama 13b Python hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 13b hf
Meta · 13B

CodeLlama 13b hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 7b Instruct hf
Meta · 6.7B

CodeLlama 7b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
CodeLlama 7b hf
Meta · 6.7B

CodeLlama 7b hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable code instruct 3b
Stabilityai · 2.8B

stable-code-instruct-3b is a 2.7B billion parameter decoder-only language model tuned from stable-code-3b. This model was trained on a mix of publicly available datasets, synthetic datasets using Direct Preference Optimization (DPO).

16K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Viking 33B
LumiOpen · 33B

Viking 33B is a 33B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1300B billion as of this release). Viking 33B is a fully open source model and is made available under the Apache 2.0 License.

4K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Viking 13B
LumiOpen · 14B

Viking 13B is a 13B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1.3 trillion as of this release). Viking 13B is a fully open source model and is made available under the Apache 2.0 License.

4K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Viking 7B
LumiOpen · 7.6B

Viking 7B is a 7B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It has been trained on 2 trillion tokens. Viking 7B is a fully open source model and is made available under the Apache 2.0 License.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
gemma 7b it
Google · 8.5B

gemma 7b it ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 7b
Google · 8.5B

gemma 7b ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2b it
Google · 2.5B

gemma 2b it ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 2b
Google · 2.5B

gemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bunny phi 2 siglip lora
BAAI

Bunny is a family of lightweight but powerful multimodal models. It offers multiple plug-and-play vision encoders, like EVA-CLIP, SigLIP and language backbones, including Phi-1.5, StableLM-2, Qwen1.5 and Phi-2. To compensate for the decrease in model size, we construct more informative training data by curated selection from a broader data source. Remarkably, our Bunny-3B model built upon SigLIP and Phi-2 outperforms the state-of-the-art MLLMs, not only in comparison with models of similar size but also against larger MLLM frameworks (7B), and even achieves performance on par with 13B models.

2K Kontext Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
bge m3 unsupervised
BAAI · 0.6B

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

8K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge m3
BAAI

For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding

8K Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
deepseek coder 7b instruct v1.5
DeepSeek · 7B

deepseek coder 7b instruct v1.5 ist ein quelloffenes Sprachmodell von DeepSeek mit 7B Parametern und einem Kontextfenster von 4K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 2 zephyr 1 6b
Stabilityai · 1.6B

Stable LM 2 Zephyr 1.6B is a 1.6 billion parameter instruction tuned language model inspired by HugginFaceH4's Zephyr 7B training pipeline. The model is trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 2 1 6b
Stabilityai · 1.6B

Please note: For commercial use, please refer to https://stability.ai/license

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek moe 16b chat
DeepSeek · 16B

python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig

4K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable code 3b
Stabilityai · 2.8B

Please note: For commercial use, please refer to https://stability.ai/license.

16K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek moe 16b base
DeepSeek · 16B

modelname = "deepseek-ai/deepseek-moe-16b-base" tokenizer = AutoTokenizer.frompretrained(modelname) model = AutoModelForCausalLM.frompretrained(modelname, torchdtype=torch.bfloat16, devicemap="auto") model.generationconfig = GenerationConfig.frompretrained(modelname) model.generationconfig.padtokenid = model.generationconfig.eostokenid

4K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
phi 2
Microsoft · 2.8B

Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-2 showcased a nearly state-of-the-art performance among models with less than 13 billion parameters.

2K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
LlamaGuard 7b
Meta · 6.7B

LlamaGuard 7b ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek llm 67b base
DeepSeek · 67B

Introducing DeepSeek LLM, an advanced language model comprising 67 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.

4K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek llm 7b chat
DeepSeek · 7B

Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek llm 7b base
DeepSeek · 7B

Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
sd turbo
Stabilityai · 0.9B

Please note: For commercial use, please refer to https://stability.ai/license.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Aquila2 70B Expr
BAAI · 70B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B, Aquila2-34B and Aquila2-70B-Expr , as well as AquilaChat2, the chat models, namely AquilaChat2-7B, AquilaChat2-34B and AquilaChat2-70B-Expr, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

4K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaChat2 70B Expr
BAAI · 70B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B, Aquila2-34B and Aquila2-70B-Expr , as well as AquilaChat2, the chat models, namely AquilaChat2-7B, AquilaChat2-34B and AquilaChat2-70B-Expr, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

4K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 72B
Qwen · 72B

通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。

33K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm zephyr 3b
Stabilityai · 2.8B

Please note: For commercial use, please refer to https://stability.ai/license.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codellama13b instruct 260k synthesis
Stabilityai · 13B

codellama13b instruct 260k synthesis ist ein quelloffenes Sprachmodell von Stabilityai mit 13B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

16K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stable diffusion xl
Stabilityai · 2.6B

japanese stable diffusion xl ist ein multimodales Sprachmodell von Stabilityai mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm instruct ja vocab beta 7b
Stabilityai · 6.9B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm base ja vocab beta 7b
Stabilityai · 6.9B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm instruct beta 70b
Stabilityai · 69B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm instruct beta 7b
Stabilityai · 6.7B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm base beta 70b
Stabilityai · 69B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm base beta 7b
Stabilityai · 6.7B

A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek coder 1.3b instruct
DeepSeek · 1.3B

Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various

16K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek coder 6.7b instruct
DeepSeek · 6.7B

Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various

16K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
JudgeLM 33B v1.0
BAAI · 33B

--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---

2K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek coder 1.3b base
DeepSeek · 1.3B

Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various

16K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
JudgeLM 13B v1.0
BAAI · 13B

--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---

2K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
JudgeLM 7B v1.0
BAAI · 7B

--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---

2K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek coder 6.7b base
DeepSeek · 6.7B

Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various

16K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Poro 34B
LumiOpen · 34B

Poro is a 34B parameter decoder-only transformer pretrained on Finnish, English and code. It was trained on 1 trillion tokens. Poro is a fully open source model and is made available under the Apache 2.0 License.

GPU ab €5,09/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
japanese stablelm instruct gamma 7b
Stabilityai · 7.2B

This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm base gamma 7b
Stabilityai · 7.2B

This is a 7B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, Mistral-7B-v0.1, to transfer the model's knowledge and capabilities to Japanese.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm 3b 4e1t instruct
Stabilityai · 2.8B

This is a 3B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese StableLM-3B-4E1T Base.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
japanese stablelm 3b 4e1t base
Stabilityai · 2.8B

This is a 3B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, StableLM-3B-4E1T, to transfer the model's knowledge and capabilities to Japanese.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaChat2 34B 16K
BAAI · 34B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

16K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaChat2 7B 16K
BAAI · 7B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

16K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Aquila2 34B
BAAI · 34B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

8K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaChat2 34B
BAAI · 34B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

4K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaChat2 7B
BAAI · 7B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

2K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Aquila2 7B
BAAI · 7.7B

We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k

8K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
llm embedder
BAAI · 0.1B

More details please refer to our Github: FlagEmbedding.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm 3b 4e1t
Stabilityai · 2.8B

StableLM-3B-4E1T is a 3 billion parameter decoder-only language model pre-trained on 1 trillion tokens of diverse English and code datasets for 4 epochs.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Mistral 7B Instruct v0.1
Mistral · 7.2B

py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Mistral 7B v0.1
Mistral · 7.2B

The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters. Mistral-7B-v0.1 outperforms Llama 2 13B on all benchmarks we tested.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
bge reranker large
BAAI · 0.6B

We have updated the new reranker, supporting larger lengths, more languages, and achieving better performance.

514 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge small zh v1.5
BAAI · 0B

More details please refer to our Github: FlagEmbedding.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge large zh v1.5
BAAI

For more details please refer to our Github: FlagEmbedding.

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
bge base zh v1.5
BAAI

More details please refer to our Github: FlagEmbedding.

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
bge small en v1.5
BAAI · 0B

More details please refer to our Github: FlagEmbedding.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge large en v1.5
BAAI · 0.3B

For more details please refer to our Github: FlagEmbedding.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge base en v1.5
BAAI · 0.1B

For more details please refer to our Github: FlagEmbedding.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
phi 1
Microsoft · 1.4B

The language model Phi-1 is a Transformer with 1.3 billion parameters, specialized for basic Python coding. Its training involved a variety of data sources, including subsets of Python codes from The Stack v1.2, Q&A content from StackOverflow, competition code from codecontests, and synthetic Python textbooks and exercises generated by gpt-3.5-turbo-0301. Even though the model and the datasets are relatively small compared to contemporary Large Language Models (LLMs), Phi-1 has demonstrated an impressive accuracy rate exceeding 50% on the simple Python coding benchmark, HumanEval.

2K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
phi 1 5
Microsoft · 1.4B

The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1, augmented with a new data source that consists of various NLP synthetic texts. When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-1.5 demonstrates a nearly state-of-the-art performance among models with less than 10 billion parameters.

2K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablecode completion alpha 3b 4k
Stabilityai · 3.3B

StableCode-Completion-Alpha-3B-4K is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that topped the stackoverflow developer survey.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablecode instruct alpha 3b
Stabilityai · 3.3B

stablecode instruct alpha 3b ist ein quelloffenes Sprachmodell von Stabilityai mit 3.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge small en
BAAI · 0B

Recommend switching to newest BAAI/bge-small-en-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge base en
BAAI · 0.1B

Recommend switching to newest BAAI/bge-base-en-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge small zh
BAAI

Recommend switching to newest BAAI/bge-small-zh-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
bge base zh
BAAI · 0.1B

Recommend switching to newest BAAI/bge-base-zh-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge large zh
BAAI · 0.3B

Recommend switching to newest BAAI/bge-large-zh-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
bge large en
BAAI · 0.3B

Recommend switching to newest BAAI/bge-large-en-v1.5, which has more reasonable similarity distribution and same method of usage.

512 Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablecode completion alpha 3b
Stabilityai · 3B

StableCode-Completion-Alpha-3B is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that were the top used languages based on the 2023 stackoverflow developer survey.

16K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
StableBeluga 13B
Stabilityai · 13B

Use Stable Chat (Research Preview) to test Stability AI's best language models for free

4K Kontext Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
StableBeluga 7B
Stabilityai · 6.7B

Use Stable Chat (Research Preview) to test Stability AI's best language models for free

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codegeex2 6b int4
Z.AI · 6B

BF16/FP16版本|BF16/FP16 version codegeex2-6b

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
StableBeluga2
Stabilityai

Use Stable Chat (Research Preview) to test Stability AI's best language models for free

4K Kontext Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
StableBeluga1 Delta
Stabilityai · 65B

Stable Beluga 1 is a Llama65B model fine-tuned on an Orca style Dataset

2K Kontext GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
AquilaCode py
BAAI

Aquila Language Model is the first open source language model that supports both Chinese and English knowledge, commercial license agreements, and compliance with domestic data regulations.

2K Kontext Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
Llama 2 70b chat hf
Meta · 69B

Llama 2 70b chat hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 2 7b chat hf
Meta · 6.7B

Llama 2 7b chat hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 2 7b hf
Meta · 6.7B

Llama 2 7b hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 2 13b hf
Meta · 13B

Llama 2 13b hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 2 13b chat hf
Meta · 13B

Llama 2 13b chat hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 2 70b hf
Meta · 69B

Llama 2 70b hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
OPI Galactica 6.7B
BAAI · 6.7B

OPI Galactica 6.7B ist ein quelloffenes Sprachmodell von BAAI mit 6.7B Parametern und einem Kontextfenster von 2K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

2K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codeexecutor
Microsoft · 0.1B

codeexecutor ist ein quelloffenes Sprachmodell von Microsoft mit 0.1B Parametern und einem Kontextfenster von 1K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

1K Kontext GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stable diffusion xl base 0.9
Stabilityai · 2.6B

stable diffusion xl base 0.9 ist ein multimodales Sprachmodell von Stabilityai mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm tuned alpha 7b
Stabilityai · 7B

StableLM-Tuned-Alpha is a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm tuned alpha 3b
Stabilityai · 3B

StableLM-Tuned-Alpha is a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm base alpha 3b
Stabilityai · 3B

📢 DISCLAIMER: The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.

4K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
stablelm base alpha 7b
Stabilityai · 7B

📢 DISCLAIMER: The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.

4K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
chatglm 6b int4 qe
Z.AI · 6B

ChatGLM-6B-INT4-QE 是 ChatGLM-6B 量化后的模型权重。具体的,ChatGLM-6B-INT4-QE 对 ChatGLM-6B 中的 28 个 GLM Block 、 Embedding 和 LM Head 进行了 INT4 量化。量化后的模型权重文件仅为 3G ,理论上 6G 显存(使用 CPU 即 6G 内存)即可推理,具有在嵌入式设备(如树莓派)上运行的可能。

GPU ab €1,68/Std. Embeddings In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
BiomedNLP BiomedELECTRA large uncased abstract
Microsoft

This model was previously named "PubMedELECTRA large (abstracts)". You can either adopt the new model name "microsoft/BiomedNLP-BiomedELECTRA-large-uncased-abstract" or update your transformers library to version 4.22+ if you need to refer to the old name.

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
whisper large v2
OpenAI · 1.5B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper medium.en
OpenAI · 0.8B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper small.en
OpenAI · 0.2B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper base.en
OpenAI · 0.1B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper tiny.en
OpenAI · 0B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper large
OpenAI · 1.5B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper medium
OpenAI · 0.8B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper small
OpenAI · 0.2B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper base
OpenAI · 0.1B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
whisper tiny
OpenAI · 0B

Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.

GPU ab €1,68/Std. Transkription In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
reacc py retriever
Microsoft

This is the retrieval model for ReACC: A Retrieval-Augmented Code Completion Framework.

514 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
unixcoder base nine
Microsoft

- Developed by: Microsoft Team - Shared by [Optional]: Hugging Face - Model type: Feature Engineering - Language(s) (NLP): en - License: Apache-2.0 - Related Models: - Parent Model: RoBERTa - Resources for more information: - Associated Paper

1K Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
unixcoder base
Microsoft

- Developed by: Microsoft Team - Shared by [Optional]: Hugging Face - Model type: Feature Engineering - Language(s) (NLP): en - License: Apache-2.0 - Related Models: - Parent Model: RoBERTa - Resources for more information: - Associated Paper

1K Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
unixcoder base unimodal
Microsoft

unixcoder base unimodal ist ein quelloffenes Sprachmodell von Microsoft mit einem Kontextfenster von 1K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

1K Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
DialoGPT medium
Microsoft

DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.

Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
muril large cased
Google

This model uses a BERT large architecture [1] pretrained from scratch using the Wikipedia [2], Common Crawl [3], PMINDIA [4] and Dakshina [5] corpora for 17 [6] Indian languages.

512 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
DialoGPT large
Microsoft

DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.

Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
CodeGPT small py
Microsoft

CodeGPT small py ist ein quelloffenes Sprachmodell von Microsoft, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.05 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
DialoGPT small
Microsoft · 0.2B

DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
codebert base
Microsoft

codebert base ist ein quelloffenes Sprachmodell von Microsoft mit einem Kontextfenster von 514 Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

514 Kontext Embeddings In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
Qwen3 30B A3B Instruct 2507
Qwen · 30B

Qwen3 30B A3B Instruct 2507 ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.25 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · warm Dedicated · verfügbar
Qwen3 4B Thinking 2507
Qwen · 4B

Qwen3 4B Thinking 2507 ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill 1.5B
DeepSeek · 1.5B

DeepSeek R1 Distill 1.5B ist ein quelloffenes Sprachmodell von DeepSeek mit 1.5B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 Coder Next
Qwen

Qwen3 Coder Next ist ein quelloffenes Sprachmodell von Qwen mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.40 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Wan 2.2 T2V 14B
Wan-AI

Open-weight video generation (text to video). We host it for you on dedicated GPUs in an EU datacenter. Contact us for a quote.

GPU ab €1,68/Std. Video In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
translategemma 4b
Google · 4B

translategemma 4b ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Super Loes (Kimi K3)
Hyai

Super Loes (Kimi K3) ist ein quelloffenes Sprachmodell von Hyai mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €3.17 pro 1M Input In der EU gehostet Router · auf Anfrage Dedicated · auf Anfrage
Stable Diffusion XL
Stability AI

Open-weight image generation (text to image). We host it for you on a dedicated GPU in an EU datacenter. Contact us for a quote.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3Guard Gen 0.6B
Qwen · 0.6B

Qwen3Guard Gen 0.6B ist ein quelloffenes Sprachmodell von Qwen mit 0.6B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 30B A3B
Qwen · 30B

Qwen3 VL 30B A3B ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 8B Thinking
Qwen · 8B

Qwen3 VL 8B Thinking ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.10 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 4B
Qwen · 4B

Qwen3 VL 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen3 VL 2B
Qwen · 2B

Qwen3 VL 2B ist ein quelloffenes Sprachmodell von Qwen mit 2B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

262K Kontext Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 2.5 Coder 1.5B
Qwen · 1.5B

Qwen 2.5 Coder 1.5B ist ein quelloffenes Sprachmodell von Qwen mit 1.5B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 3B
Qwen · 3B

Qwen2.5 VL 3B ist ein quelloffenes Sprachmodell von Qwen mit 3B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
deepseek coder 6.7b
DeepSeek · 6.7B

deepseek coder 6.7b ist ein quelloffenes Sprachmodell von DeepSeek mit 6.7B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

16K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3n E2B
Google · 2B

gemma 3n E2B ist ein quelloffenes Sprachmodell von Google mit 2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Gemma 3 4B
Google · 4B

Gemma 3 4B ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Gemma 3 1B
Google · 1B

Gemma 3 1B ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
gemma 3 12b
Google · 12B

gemma 3 12b ist ein quelloffenes Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
functiongemma 270m
Google

functiongemma 270m ist ein quelloffenes Sprachmodell von Google, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.02 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
FLUX.1 schnell
Black Forest Labs

Open-weight image generation (text to image), served on a dedicated GPU in an EU datacenter via an OpenAI-compatible /v1/images/generations endpoint. Deploy it as a dedicated instance from the wizard, or contact us for help sizing it.

GPU ab €1,68/Std. Bild In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek Coder V2 Lite (16B)
DeepSeek · 16B

DeepSeek Coder V2 Lite (16B) ist ein quelloffenes Sprachmodell von DeepSeek mit 16B Parametern und einem Kontextfenster von 164K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

164K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek V3 (685B MoE)
DeepSeek · 685B

DeepSeek V3 (685B MoE) ist ein quelloffenes Sprachmodell von DeepSeek mit 685B Parametern und einem Kontextfenster von 164K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

164K Kontext Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Kokoro 82M (TTS)
Hexgrad

Open-weight text-to-speech. We host it for you on a dedicated GPU in an EU datacenter, served via an OpenAI-compatible /v1/audio/speech endpoint. Contact us for a quote.

GPU ab €1,68/Std. Sprachausgabe In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 VL 32B
Qwen · 32B

Qwen2.5 VL 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

128K Kontext GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 2.5 14B
Qwen · 14B

Qwen 2.5 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 Coder 14B
Qwen · 14B

Qwen2.5 Coder 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen2.5 7B
Qwen · 7B

Qwen2.5 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 3 4B
Qwen · 4B

Qwen 3 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

41K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 2.5 Coder 7B
Qwen · 7B

Qwen 2.5 Coder 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.05 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 2.5 Coder 32B
Qwen · 32B

Qwen 2.5 Coder 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
DeepSeek R1 Distill 32B
DeepSeek · 32B

DeepSeek R1 Distill 32B ist ein quelloffenes Sprachmodell von DeepSeek mit 32B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Qwen 2.5 72B
Qwen · 72B

Qwen 2.5 72B ist ein quelloffenes Sprachmodell von Qwen mit 72B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

33K Kontext Router €0.40 pro 1M Input GPU ab €5,09/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 Mini (3.8B)
Microsoft · 3.8B

Phi 4 Mini (3.8B) ist ein quelloffenes Sprachmodell von Microsoft mit 3.8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

131K Kontext Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 11B Vision
Meta · 11B

Llama 3.2 11B Vision ist ein quelloffenes Sprachmodell von Meta mit 11B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.06 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Phi 4 (14B)
Microsoft · 14B

Phi 4 (14B) ist ein quelloffenes Sprachmodell von Microsoft mit 14B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

16K Kontext Router €0.08 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
medgemma 4b
Google · 4B

medgemma 4b ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.03 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
medgemma 27b
Google · 27B

medgemma 27b ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.15 pro 1M Input GPU ab €1,68/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Loes
HostYourAI · 7B

Sovereign EU model fine-tuned by HostYourAI on dutch-clean.

Router €0.05 pro 1M Input GPU ab €1,68/Std. 🇪🇺 Europäisch Router · auf Anfrage Dedicated · verfügbar
Llama 4 Maverick (17Bx128E)
Meta · 17B

Llama 4 Maverick (17Bx128E) ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €0.40 pro 1M Input GPU ab €54,00/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 3.2 90B Vision
Meta · 90B

Llama 3.2 90B Vision ist ein quelloffenes Sprachmodell von Meta mit 90B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Llama 4 Scout (17Bx16E)
Meta · 17B

Llama 4 Scout (17Bx16E) ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.

Router €1.20 pro 1M Input GPU ab €9,13/Std. In der EU gehostet Router · auf Anfrage Dedicated · verfügbar
Keine Modelle gefunden. Passen Sie Ihre Suche oder Filter an.

Hosten. Routen. Ausliefern.

Keine Kreditkarte nötig. Pay as you go, jederzeit kündbar.

Heute kostenlos starten