715 Open-Source-Modelle, gehostet auf GPUs in der EU. Ein OpenAI-kompatibler API-Key, Scale-to-Zero oder dediziert.
715 Modelle
Kimi K3 is an open-weight, native multimodal agentic model and our most capable model to date. It is a 2.8T-parameter model built on Kimi Delta Attention (KDA) and Attention Residuals (AttnRes), with native vision capabilities and a 1-million-token context window. It is the world's first open 3T-class model, designed for frontier intelligence across long-horizon coding, knowledge work, and reasoning.
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.
Qwen3 Embedding 8B ist ein quelloffenes Sprachmodell von Qwen mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.3 70B ist ein quelloffenes Sprachmodell von Meta mit 70B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
Llama 3.3 70B Instruct ist ein quelloffenes Sprachmodell von Llama-3.3-70b-instruct mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency - Improved Architecture: We propose IndexShare, which reuses the same indexer across every fou
Qwen3 Coder 30B A3B ist ein quelloffenes Sprachmodell von Qwen3-coder-30b-a3b-instruct mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 235B A22B Instruct ist ein quelloffenes Sprachmodell von Qwen3-235b-a22b-instruct-2507 mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
BGE Multilingual Gemma2 ist ein quelloffenes Sprachmodell von BAAI mit einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Loes Large World: Qwen3.5-27B dense base met volledig-Europese SFT-adapter (LoRA), geserveerd via vLLM --enable-lora + qwen3 reasoning-parser, CUDA-graphs. Topmodel-spoor voor chat.loes.ai.
We present a preview version of DeepSeek-V4 series, including two strong Mixture-of-Experts (MoE) language models — DeepSeek-V4-Pro with 1.6T parameters (49B activated) and DeepSeek-V4-Flash with 284B parameters (13B activated) — both supporting a context length of one million tokens.
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
Qwen2.5 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Mistral Medium 3.5 ist ein quelloffenes Sprachmodell von Mistral-medium-3.5-128b mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Mistral Nemo 12B ist ein quelloffenes Sprachmodell von Mistral mit 12B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Mistral Small 3.2 24B ist ein quelloffenes Sprachmodell von Mistral-small-3.2-24b-instruct-2506 mit einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
GPT-OSS 120B ist ein quelloffenes Sprachmodell von Gpt-oss-120b mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Gemma 4 26B A4B ist ein quelloffenes Sprachmodell von Gemma-4-26b-a4b-it mit einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Gemma 3 27B ist ein quelloffenes Sprachmodell von Gemma-3-27b-it mit einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
Qwen2.5 VL 72B ist ein quelloffenes Sprachmodell von Qwen mit 72B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 8B ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Holo2 30B A3B ist ein quelloffenes Sprachmodell von Holo2-30b-a3b mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
Qwen3.6 35B A3B ist ein quelloffenes Sprachmodell von Qwen mit 35B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Pixtral 12B ist ein quelloffenes Sprachmodell von Pixtral-12b-2409 mit einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro.
Qwen2.5 VL 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).
MiniMax M3 ist ein quelloffenes Sprachmodell von MiniMaxAI mit einem Kontextfenster von 205K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3.5 397B A17B ist ein quelloffenes Sprachmodell von Qwen3.5-397b-a17b mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Whisper Large v3 ist ein quelloffenes Sprachmodell von OpenAI, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
MiniMax M2.5 ist ein quelloffenes Sprachmodell von MiniMaxAI mit einem Kontextfenster von 197K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:
Qwen 3 0.6B ist ein quelloffenes Sprachmodell von Qwen mit 0.6B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Sovereign EU model fine-tuned by HostYourAI on loes-large-v1.
Devstral 2 123B ist ein quelloffenes Sprachmodell von Devstral-2-123b-instruct-2512 mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.
Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.
Qwen 3 Coder 30B-A3B (MoE) ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3.6 27B ist ein quelloffenes Sprachmodell von Qwen mit 27B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
This is the model card for EuroLLM-22B-Instruct. You can also check the pre-trained version: EuroLLM-22B-2515.
py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest
DeepSeek R1 Distill 14B ist ein quelloffenes Sprachmodell von DeepSeek mit 14B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 1B ist ein quelloffenes Sprachmodell von Meta mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.
Voxtral Small 24B ist ein quelloffenes Sprachmodell von Voxtral-small-24b-2507 mit einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek R1 0528 Qwen3 8B ist ein quelloffenes Sprachmodell von DeepSeek mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, TokenSpeed, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
DeepSeek-V4-Pro-0813 is the official release of DeepSeek-V4-Pro, superseding the preview version, with greatly enhanced agentic capabilities and performance improvements that are especially pronounced in production environments. It is built on the DeepSeek-V4-Pro (Preview) model structure, with a DSpark speculative decoding module attached.
The pre-training data has a cutoff date of September 2025. The post-training data has a cutoff date of May 2026.
AREX is a family of deep research agents developed by the Beijing Academy of Artificial Intelligence (BAAI). It is designed for long-horizon tasks in which an agent must search across sources, assemble candidate answers, verify multiple constraints, and revise its research plan when the available evidence is incomplete.
AREX is a family of deep research agents developed by the Beijing Academy of Artificial Intelligence (BAAI). It is designed for long-horizon tasks in which an agent must search across sources, assemble candidate answers, verify multiple constraints, and revise its research plan when the available evidence is incomplete.
Fara1.5-27B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.
Fara1.5-4B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.
GOVERNING TERMS: Use of this model is governed by the OpenMDW License Agreement, version 1.1. ADDITIONAL INFORMATION: Apache License, Version 2.0.
GOVERNING TERMS: Use of this model is governed by the OpenMDW License Agreement, version 1.1. ADDITIONAL INFORMATION: Apache License, Version 2.0.
salamandra 7b fc 2607 ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
ALIA 40b fc 2607 ist ein quelloffenes Sprachmodell von BSC-LT mit 40B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Cosmos3-Super-Text2Image-4Step is a 4-step distilled version of the base Cosmos3-Super-Text2Image model. Given a text prompt, it generates a high-fidelity image.
The NVIDIA Qwen-Image-Flash model generates images from text prompts using a four-step, DMD2-distilled version of Qwen/Qwen-Image. The distillation used DMD2 from NVIDIA FastGen, NVIDIA Model Optimizer, and NVIDIA AutoModel while retaining the base model architecture. The packaged scheduler is configured for the four-step, shift-3 trajectory.
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
GELab Zero 4B preview Sico Evolution ist ein multimodales Sprachmodell von Microsoft mit 4.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. The 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs.
The model employs a hybrid MoE architecture with interleaved Mamba, MoE, and Attention layers. Like Nemotron-3-Super, it supports Multi-Token Prediction (MTP) for faster text generation. Compared to its parent, Puzzle-75B-A9B reduces the model from 120.7B total / 12.8B active parameters to 75.3B total / 9.3B active parameters.
Poro 2 Long Instruct is an instruction-following chatbot model with extended context support, created through supervised fine-tuning (SFT) of the Poro 2 Long Base model followed by merging the SFT checkpoint back with the base model to preserve long-context performance. This model is designed for conversational AI applications and instruction following in both Finnish and English, with support for context lengths up to 128K tokens. It was trained on a carefully curated mix of English and Finnish instruction data.
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
For more details on how to deploy and use the model - see the Quick Start Guide below!
For more details on how to deploy and use the model - see the Quick Start Guide below!
Cosmos3 is a collection of Omnimodal world models capable of generating dynamic, high-quality video, image, audio, and action commands from combinations of text, image, video, and action trajectory inputs. It serves as a foundational building block for a broad range of Physical AI applications and research spanning world understanding, world generation, simulation, and embodied policy learning.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
MagenticBrain is a 14B-parameter orchestration model from Microsoft Research AI Frontiers. It plans multi-step tasks, calls declared tools, and coordinates sub-agents. It does not execute actions itself — every real-world side effect happens inside a host harness.
Fara1.5-9B is a multimodal computer use agent (CUA) for web browsers, from Microsoft Research AI Frontiers. It observes the browser through screenshots and acts on the user's behalf by emitting structured tool calls — click, type, scroll, visit URL, web search, and so on — to complete tasks end-to-end.
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
Poro 2 8B Math Reasoning RL Preview is a specialized model focused on mathematical reasoning and problem-solving. This preview model was created through reinforcement learning (RL) on top of the Math Reasoning SFT checkpoint. This model excels at mathematical reasoning tasks but is not optimized for general conversational use or other domains.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
NVIDIA Cosmos Reason 2 is an open, customizable, 32B-parameter reasoning vision language model (VLM) for physical AI and robotics that enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
Model Summary: Granite-Embedding-311M-Multilingual-R2 is a 311M parameter dense embedding model from the Granite Embeddings collection for high-quality multilingual text embeddings. It produces 768-dimensional vectors with a context length of up to 32,768 tokens. The model supports 200+ languages (based on the multilingual pretraining corpus of the underlying encoder), with enhanced support for 52 languages and programming code that receive explicit retrieval-pair and cross-lingual training. All training data uses permissive, enterprise-friendly licenses, plus IBM-collected and IBM-generated d
Model Summary: Granite-Embedding-97M-Multilingual-R2 is a 97M parameter dense embedding model from the Granite Embeddings collection for high-quality multilingual text embeddings at minimal compute cost. It produces 384-dimensional vectors with a context length of up to 32,768 tokens. The model supports 200+ languages (based on the multilingual pretraining corpus of the underlying encoder), with enhanced support for 52 languages and programming code that receive explicit retrieval-pair and cross-lingual training. All training data uses permissive, enterprise-friendly licenses, plus IBM-collect
Granite Guardian 4.1 8B introduces improved Bring Your Own Criteria (BYOC) support, enabling users to define arbitrary judging criteria beyond the pre-baked safety and hallucination detectors. The model can now faithfully evaluate complex, multi-part requirements such as formatting rules, length constraints, and domain-specific instructions.
Model Summary: Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint, providing a lightweight alternative to much larger frontier models for these tasks:
Granite-Speech-4.1-2B-Plus has similar capabilities to the Granite-Speech-4.1-2B model. The plus model adds two new community-requested rich transcription features that can be activated with a simple prompt change: speaker-attributed ASR (speaker labels and word transcripts) and word-level timing information. Unlike the base mode, the plus model doesn't provide punctuation and capitalization.
Model Summary: Granite Speech 4.1 2B is a compact and efficient speech-language model, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST) for English, French, German, Spanish, Portuguese and Japanese.
Poro 2 Long Base is an 8B parameter decoder-only transformer created by extending Poro 2 8B Base from an 8K to 128K token context window using LongRoPE. The model supports both English and Finnish with an extended context window of 128K tokens. Poro 2 Long Base is released as a fully open source model under the Llama 3.1 Community License.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.
Model Summary: Granite-4.1-30B is a 30B parameter long-context instruct model finetuned from Granite-4.1-30B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
Model Summary: Granite‑4.1‑8B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of text‑to‑text generation tasks. In addition to standard generation, it supports Fill‑in‑the‑Middle (FIM) code completion through specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens using a five‑phase training strategy: 10 trillion tokens in phase one, 2 trillion tokens each in phases two and three, and 0.5 trillion tokens in phase four. In the final phase, long‑context extension is applied to expand the m
Model Summary: Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from Granite-4.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture and weights as Granite 4.0 3B Micro, which is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the third, and 0.5 trillion in the final stage. An additional training phase is applied
Model Summary: Granite-4.1-3B is a 3B parameter long-context instruct model finetuned from Granite-4.1-3B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
EGM-Qwen3-VL-4B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-4B-Thinking.
EGM-Qwen3-VL-8B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-8B-Thinking.
EGM-Qwen3-VL-4B is an efficient visual grounding model from the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-4B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).
Poro 2 8B Math Reasoning SFT Preview is a specialized model focused on mathematical reasoning and problem-solving. This preview model was created through supervised fine-tuning of a context-extended base model. This model excels at mathematical reasoning tasks but is not optimized for general conversational use or other domains.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
This model is a fine-tuned derivative of Qwen3.5-35B-A3B. Follow the Qwen3.5-35B-A3B serving guide for deployment with vLLM, replacing the model path with nvidia/NVIDIA-Ising-Calibration-1-35B-A3B. Suggested inference settings: temperature=0.2, maxtokens=16384.
harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.
harrier-oss-v1 is a family of multilingual text embedding models developed by Microsoft. The models use decoder-only architectures with last-token pooling and L2 normalization to produce dense text embeddings. They can be applied to a wide range of tasks, including but not limited to retrieval, clustering, semantic similarity, classification, bitext mining, and reranking. The models achieve state-of-the-art results on the Multilingual MTEB v2 benchmark as of the release date.
We introduce UniRG-CXR, a radiology report generation model that obtains SOTA performance on ReXrank. More details can be found in the paper: Scaling medical imaging report generation with multimodal reinforcement learning
gemma 4 31B ist ein quelloffenes Sprachmodell von Google mit 31B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.
Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.
The pretraining data has a cutoff date of September 2024\.
Model Dates: Trained between Oct 2025 and March 2026
Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:
EGM-Qwen3-VL-8B is the flagship model of the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-8B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).
[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger
Qwen3.5 0.8B ist ein quelloffenes Sprachmodell von Qwen mit 0.8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3.5 2B ist ein quelloffenes Sprachmodell von Qwen mit 2B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Model Summary: Granite-4.0-1b-speech is a compact and efficient speech-language model, specifically designed for multilingual automatic speech recognition (ASR) and bidirectional automatic speech translation (AST).
Qwen3.5 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger
[!Note] This repository contains model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
Qwen3.5 35B A3B ist ein quelloffenes Sprachmodell von Qwen mit 35B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
llama-nv-embed-reasoning-3b is a 3.2B-parameter embedding model designed to produce high‑quality sentence and document representations for retrieval, semantic search, and similarity tasks, with a strong focus on reasoning‑heavy content.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
We introduce X-Reasoner, a vision-language model posttrained solely on general-domain text for generalizable reasoning, using a twostage approach: an initial supervised fine-tuning phase with distilled long chainof-thoughts, followed by reinforcement learning with verifiable rewards. Experiments show that X-Reasoner successfully transfers reasoning capabilities to both multimodal and out-of-domain settings, outperforming existing state-of-theart models trained with in-domain and multimodal data across various general and medical benchmarks. More details can be found in the paper: X-Reasoner: T
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, GLM-OCR deliver
Chart2CSV is a specialized vision-language model fine-tuned for the accurate extraction of tabular data from charts and visualizations. Built on top of ibm-granite/granite-vision-3.3-2b, it produces machine-readable CSV outputs with improved numeric fidelity compared to general-purpose VLMs. The model is trained using code-guided synthetic chart data following the ChartGen methodology, which strengthens factual grounding and reduces hallucination in the Chart-to-CSV task.
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:
The Qwen3-ASR family includes Qwen3-ASR-1.7B and Qwen3-ASR-0.6B, which support language identification and ASR for 52 languages and dialects. Both leverage large-scale speech training data and the strong audio understanding capability of their foundation model, Qwen3-Omni. Experiments show that the 1.7B version achieves state-of-the-art performance among open-source ASR models and is competitive with the strongest proprietary commercial APIs. Here are the main features:
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:
This is the model card for EuroLLM-9B-Instruct-2512, an improved version of utter-project/EuroLLM-9B-Instruct. In comparison with the previous version, this version includes the long-context extension phase and the revamped post-training recipe from utter-project/EuroLLM-22B-Instruct.
Developer: Microsoft Corporation Authorized Representative: Microsoft Ireland Operations Limited, 70 Sir John Rogerson's Quay, Dublin 2, D02 R296, Ireland Release Date: March 4, 2026 License: MIT Parameters: 15B Context Length: 16,384 tokens Inputs: Text and Images Outputs: Text Training GPUs: 240 B200s Training Time: 4 days Training Dates: February 3, 2025 – February 21, 2026 Model Dependencies: Phi-4-Reasoning
GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.
translategemma 27b it ist ein multimodales Sprachmodell von Google mit 29B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
translategemma 12b it ist ein multimodales Sprachmodell von Google mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
translategemma 4b it ist ein multimodales Sprachmodell von Google mit 5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
GLM-Image is an image generation model adopts a hybrid autoregressive + diffusion decoder architecture. In general image generation quality, GLM‑Image aligns with mainstream latent diffusion approaches, but it shows significant advantages in text-rendering and knowledge‑intensive generation scenarios. It performs especially well in tasks requiring precise semantic understanding and complex information expression, while maintaining strong capabilities in high‑fidelity and fine‑grained detail generation. In addition to text‑to‑image generation, GLM‑Image also supports a rich set of image‑to‑imag
This model is a fine-tuned version of the openai/whisper-large-v3-turbo model finetuned for automatic speech recognition (ASR) in several Kenyan languages, including Swahili, Kalenjin, Kikuyu, Luo, Maasai and Somali. Whisper is a transformer-based encoder-decoder model that converts raw audio into text. The encoder processes audio inputs as log-Mel spectrograms, capturing acoustic and linguistic features, while the decoder generates text tokens in an autoregressive manner. This design allows the model to handle diverse languages, accents, and noise conditions with strong generalization.
medgemma 1.5 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Fine-tuning was performed on the entire unified multilingual ASR dataset, comprising the mentioned six languages, to encourage cross-lingual generalization. During fine-tuning, only the audio-specific components: audio embedding module, audio encoder, and audio projection layers, were unfrozen and set as trainable, while the rest of the model parameters remained frozen to preserve pretrained language capabilities. Dropout was applied to both the audio encoder and projection layers to regularize training. The model leverages a multimodal processor that handles text tokenization and audio featur
The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.
The Qwen3-VL-Embedding and Qwen3-VL-Reranker model series are the latest additions to the Qwen family, built upon the recently open-sourced and powerful Qwen3-VL foundation model. Specifically designed for multimodal information retrieval and cross-modal understanding, this suite accepts diverse inputs including text, images, screenshots, and videos, as well as inputs containing a mixture of these modalities.
We are excited to introduce Qwen-Image-2512, the December update of Qwen-Image’s text-to-image foundational model. You are welcome to try the latest model at Qwen Chat. Compared to the base Qwen-Image model released in August, Qwen-Image-2512 features the following key improvements:
GLM-4.7, your new coding partner, is coming with the following features:
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
This is the model card for EuroMoE-2.6B-A0.6B-Instruct-2512. You can also check the pre-trained version: EuroMoE-2.6B-A0.6B-2512.
Cosmos Reason2 2B ist ein multimodales Sprachmodell von NVIDIA mit 2.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Cosmos Reason2 8B ist ein multimodales Sprachmodell von NVIDIA mit 8.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
GLM-ASR-Nano-2512 is a robust, open-source speech recognition model with 1.5B parameters. Designed for real-world complexity, it outperforms OpenAI Whisper V3 on multiple benchmarks while maintaining a compact size.
⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.
⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
llama-nemotron-embed-vl-1b-v2 was developed by NVIDIA for multimodal question-answering retrieval. The model can embed document pages in the form of image, text, or combined image–text inputs. Documents can be retrieved given a user query in text form. The model supports page images containing text, tables, charts, and infographics. We report the evaluation of this model on two internal multimodal retrieval benchmarks, and on the popular ViDoRe V1 and V2 benchmarks and the new Vidore V3 benchmark.
This is the model card for EuroLLM-22B. You can also check the post-trained version: EuroLLM-22B-Instruct-2515.
[!Tip] This model was contributed by Xenova from Hugging Face. We sincerely appreciate the integration and community collaboration. While preliminary functionality checks have been performed, comprehensive testing has not yet been completed. We recommend you to proceed with caution and conducting your own evaluations for specific use cases. If any issues arise, open a PR/Issue here and we will try to address them promptly.
Salamandra-VL-7B-2512 is the latest version of the Salamandra vision model family. This version brings significant improvements in both architecture and training data.
Voxtral TTS is a frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents. The model is released with BF16 weights and a set of reference voices. These voices are licensed under CC BY-NC 4, which is the license that the model inherits.
- Repository: https://github.com/zheny2751-dotcom/WebVIA - Paper: https://arxiv.org/pdf/2511.06251
- Repository: https://github.com/zai-org/UI2CodeN - Paper: https://arxiv.org/abs/2511.08195
Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage.
Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.
Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.
Update: We just released Fara1.5 which improves on Fara-7B dramatically and is available in three model sizes 4B, 9B and 27B!
Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.
Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.
This model is for research and development only. <br
- Repository: https://github.com/thu-coai/Glyph - Paper: https://arxiv.org/abs/2510.17800
Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.
Using the 🤗's Diffusers library to run URSA in a simple and efficient manner.
NVIDIA-Nemotron-Nano-VL-12B-V2-FP4-QAD is the quantized version of the NVIDIA Nemotron Nano VL V2 model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Nemotron Nano VL FP4 QAD model is quantized with TensorRT Model Optimizer.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation
The Llama Nemotron Embedding 1B model is optimized for multilingual and cross-lingual text question-answering retrieval with support for long documents (up to 8192 tokens) and dynamic embedding size (Matryoshka Embeddings). This model was evaluated on 26 languages: English, Arabic, Bengali, Chinese, Czech, Danish, Dutch, Finnish, French, German, Hebrew, Hindi, Hungarian, Indonesian, Italian, Japanese, Korean, Norwegian, Persian, Polish, Portuguese, Russian, Spanish, Swedish, Thai, and Turkish.
This model is for research and development only.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
functiongemma 270m it ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Model Summary: Granite-4.0-350M-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the
Model Summary: Granite-4.0-350M is a lightweight instruct model finetuned from Granite-4.0-350M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the th
Model Summary: Granite-4.0-1B is a lightweight instruct model finetuned from Granite-4.0-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
This model achieves state-of-the-art performance on the multilingual MTEB leaderboard as of October 21, 2025.
Llama-3.1-Nemotron-Nano-VL-8B-V1-FP4-QAD is the quantized version of the NVIDIA Llama Nemotron Nano VL model, which is an auto-regressive vision language model that uses an optimized transformer architecture. For more information, please check here. The NVIDIA Llama Nemotron Nano VL FP4 QAD model is quantized with TensorRT Model Optimizer.
Unlike typical LLMs that are trained to play the role of the "assistant" in conversation, we trained UserLM-8b to simulate the “user” role in conversation (by training it to predict user turns in a large corpus of conversations called WildChat). This model is useful in simulating more realistic conversations, which is in turn useful in the development of more robust assistants.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Compared with GLM-4.5, GLM-4.6 brings several key improvements:
[!WARNING] WARNING: This is a language model that has undergone instruction tuning for conversational settings that exploit function calling capabilities. It has not been aligned with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications. How to use from datetime import datetime from transformers import AutoTokenizer, AutoModelForCausalLM import transformer
We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.
This update maintains the model's original capabilities while addressing issues reported by users, including:
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2. It is a state-of-the-art mixture-of-experts (MoE) language model, featuring 32 billion activated parameters and a total of 1 trillion parameters.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
embeddinggemma 300m qat q8 0 unquantized ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
embeddinggemma 300m qat q4 0 unquantized ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:
TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
The pretraining data has a cutoff date of September 2024.
TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
gemma 3 270m ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 4B Instruct 2507 ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
Install the latest version of diffusers pip install git+https://github.com/huggingface/diffusers
Llama-3.3-Nemotron-Super-49B-v1.5-FP8 is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and agentic tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
gemma 3 270m it ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:
We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:
Compared to nvidia/DeepSeek-R1-0528-FP4, this checkpoint additionally quantizes the wo module in attention layers.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
We are excited to introduce Wan2.2, a major upgrade to our foundational video models. With Wan2.2, we have focused on incorporating the following innovations:
Model Summary: Granite-embedding-small-english-r2 is a 47M parameter dense biencoder embedding model from the Granite Embeddings collection that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 based on context length of upto 8192 tokens. Compared to most other open-source models, this model was only trained using open-source relevance-pair datasets with permissive, enterprise-friendly license, plus IBM collected and generated datasets.
Model Summary: Granite-embedding-english-r2 is a 149M parameter dense biencoder embedding model from the Granite Embeddings collection that can be used to generate high quality text embeddings. This model produces embedding vectors of size 768 based on context length of upto 8192 tokens. Compared to most other open-source models, this model was only trained using open-source relevance-pair datasets with permissive, enterprise-friendly license, plus IBM collected and generated datasets.
embeddinggemma 300m ist ein quelloffenes Sprachmodell von Google mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
The MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e. PubMed commercial, Medical Wikipedia, Medical Guidelines, Medical Coding, and open-source clinical documents) and merged back with the SLERP method in their base model to conserve general abilities. One model combined all five experts into one general expert with the multi-model merging method BreadCrumbs. F
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
medgemma 27b it ist ein multimodales Sprachmodell von Google mit 29B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
This model was converted to MLX format from ibm-granite/granite-docling-258M using mlx-vlm version 0.3.3. Refer to the original model card for more details on the model.
FLUX.1 Krea dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Dayhoff is an Atlas of both protein sequence data and generative language models — a centralized resource that brings together 3.34 billion protein sequences across 1.7 billion clusters of metagenomic and natural protein sequences (GigaRef), 46 million structure-derived synthetic sequences (BackboneRef), and 16 million multiple sequence alignments (OpenProteinSet). These models can natively predict zero-shot mutation effects on fitness, scaffold structural motifs by conditioning on evolutionary or structural context, and perform guided generation of novel proteins within specified families. Le
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.
Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.
Cosmos Predict2 0.6B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 0.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a larger variant, Phi-mini-MoE, with 7.6B total and 2.4B activated pa
Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a smaller variant, Phi-tiny-MoE, with 3.8B total and 1.1B activated p
[!Note] This is an improved version of Kimi-VL-A3B-Thinking. Please consider using this updated model instead of the previous version.
Phi-4-mini-flash-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 64K token context length.
We introduce Kimi-Dev-72B, our new open-source coding LLM for software engineering tasks. Kimi-Dev-72B achieves a new state-of-the-art on SWE-bench Verified among open-source models.
Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 70B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.
Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 8B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.
gemma 3n E2B it ist ein multimodales Sprachmodell von Google mit 5.4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
⚠️ PREVIEW RELEASE: This is a preview version of EuroMoE-2.6B-A0.6B-Instruct-Preview. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroLLM-2.6B-A0.6-2512, the pre-trained model for EuroLLM-2.6B-A0.6-2512-Instruct.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-1.7B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-9B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroLLM-9B-2512, an improved version of utter-project/EuroLLM-9B. In comparison with the previous version, this version includes the long-context extension phase from utter-project/EuroLLM-22B.
gemma 3n E4B it ist ein multimodales Sprachmodell von Google mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code re
The Qwen3 Embedding model series is the latest proprietary model of the Qwen family, specifically designed for text embedding and ranking tasks. Building upon the dense foundational models of the Qwen3 series, it provides a comprehensive range of text embeddings and reranking models in various sizes (0.6B, 4B, and 8B). This series inherits the exceptional multilingual capabilities, long-text understanding, and reasoning skills of its foundational model. The Qwen3 Embedding series represents significant advancements in multiple text embedding and ranking tasks, including text retrieval, code re
This model was introduced in the paper GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. It is developed based on UI-TARS-2B-SFT and is designed to predict the correctness of an action position given a language instruction. This model is well-suited for GUI-Actor, as its attention map effectively provides diverse candidates for verification with only a single inference.
Poro 2 70B Base is a 70B parameter decoder-only transformer created through continued pretraining of Llama 3.1 70B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.
Poro 2 8B Base is an 8B parameter decoder-only transformer created through continued pretraining of Llama 3.1 8B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.
Poro 2 70B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 70B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.
Poro 2 8B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 8B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.
[!WARNING] WARNING: This model has been deprecated and is no longer recommended. For the latest model, please visit: https://huggingface.co/BSC-LT/Salamandra-VL-7B-2512
medgemma 27b text it ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
medgemma 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Granite Docling 258M builds upon the Idefics3 architecture, but introduces two key modifications: it replaces the vision encoder with siglip2-base-patch16-512 and substitutes the language model with a Granite 165M LLM. Try out our Granite-Docling-258 demo today.
2025-4-2 🌟🌟 BGE-VL models are also available on WiseModel.
2025-04-06 🚀🚀 MVRB Dataset are released on Huggingface: MVRB
For more details please refer to our Github: FlagEmbedding.
2025-4-2 🌟🌟 BGE-VL models are also available on WiseModel.
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
- Model Description - Intended Uses and Limitations - How to Get Started with the Model - Training Details - Citation - Additional Information
NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits (ICML'2025)
Model Summary: Granite-4-Tiny-Preview is a 7B parameter fine-grained hybrid mixture-of-experts (MoE) instruct model fine-tuned from Granite-4.0-Tiny-Base-Preview using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised fine-tuning, and model alignment using reinforcement learning.
Phi-4-mini-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 128K token context length.
Model Summary: Granite-speech-3.3-2b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.3-2b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.3-2b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 30B A3B ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 8B ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 1.7B ist ein quelloffenes Sprachmodell von Qwen mit 1.7B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
We present Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation. This repository hosts the model checkpoints for Kimi-Audio-7B.
We present Kimi-Audio, an open-source audio foundation model excelling in audio understanding, generation, and conversation. This repository hosts the model checkpoints for Kimi-Audio-7B-Instruct.
Llama Guard 4 12B ist ein multimodales Sprachmodell von Meta mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Cosmos Predict2 14B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 14B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Cosmos Predict2 2B Text2Image ist ein multimodales Sprachmodell von NVIDIA mit 2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
NVIDIA Cosmos Reason – an open, customizable, 7B-parameter reasoning vision language model (VLM) for physical AI and robotics - enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.
[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).
Model Summary: Granite-speech-3.3-8b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.3-8b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.3-8b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i
[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).
Model Summary: Granite-3.3-8B-Instruct is a 8-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-8B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on
Model Summary: Granite-3.3-2B-Instruct is a 2-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-2B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on
[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities—all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B).
gemma 3 12b it qat q4 0 unquantized ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
Llama 4 Maverick 17B 128E ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 4 Scout 17B 16E ist ein multimodales Sprachmodell von Meta mit 109B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 4 Scout 17B 16E Instruct ist ein multimodales Sprachmodell von Meta mit 109B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 4 Maverick 17B 128E Instruct ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Model Summary: Granite-speech-3.2-8b is a compact and efficient speech-language model, specifically designed for automatic speech recognition (ASR) and automatic speech translation (AST). Granite-speech-3.2-8b uses a two-pass design, unlike integrated models that combine speech and language into a single pass. Initial calls to granite-speech-3.2-8b will transcribe audio files into text. To process the transcribed text using the underlying Granite language model, users must make a second call as each step must be explicitly initiated.
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.
txgemma 2b predict ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
salamandra 7b instruct tools ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama-3.1-Nemotron-Nano-8B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.1-8B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling.
Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.
gemma 3 1b it ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.
2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.
+ Resolution: Width and height must be between 512px and 2048px, divisible by 32, and ensure the maximum number of pixels does not exceed 2^21 px. + Precision: BF16 / FP32 (FP16 is not supported as it will cause overflow resulting in completely black images)
gemma 3 12b it ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 12b pt ist ein multimodales Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 27b it ist ein multimodales Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 27b pt ist ein multimodales Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p
In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p
In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p
In this repository, we present Wan2.1, a comprehensive and open suite of video foundation models that pushes the boundaries of video generation. Wan2.1 offers these key features: - 👍 SOTA Performance: Wan2.1 consistently outperforms existing open-source models and state-of-the-art commercial solutions across multiple benchmarks. - 👍 Supports Consumer-grade GPUs: The T2V-1.3B model requires only 8.19 GB VRAM, making it compatible with almost all consumer-grade GPUs. It can generate a 5-second 480P video on an RTX 4090 in about 4 minutes (without optimization techniques like quantization). Its p
2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.
2024-12-27 🚀🚀 BGE-VL-CLIP models are released on Huggingface: BGE-VL-base and BGE-VL-large.
🎉Phi-4: [mini-reasoning | reasoning] | [multimodal-instruct | onnx]; [mini-instruct | onnx]
- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates
- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates
gemma 3 1b pt ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 4b it ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 4b pt ist ein multimodales Sprachmodell von Google mit 4.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
🎉Phi-4: [mini-reasoning | reasoning] | [multimodal-instruct | onnx]; [mini-instruct | onnx]
Model Summary: Granite-Embedding-30m-Sparse is a 30M parameter sparse biencoder embedding model from the Granite Experimental suite that can be used to generate high quality text embeddings. This model produces variable length bag-of-word like dictionary, containing expansions of sentence tokens and their corresponding weights and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. While maintaining competitive scores on academic benchmarks such as BEIR, this model also performs well on ma
Model Summary: Granite-3.2-2B-Instruct is an 2-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-2B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic data designed for reasoning tasks. The model allows controllability of its thinking capability, ensuring it is applied only when required.
Model Summary: granite-vision-3.2-2b is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model with both image and text modalities.
Model Summary: granite-vision-3.1-2b-preview is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model (https://huggingface.co/ibm-granite/granite-3.1-2b-instruct) with
--- license: other licensename: qwen licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-72B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
--- licensename: qwen-research licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
DeepSeek R1 Distill Llama 70B ist ein quelloffenes Sprachmodell von DeepSeek mit 70B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek R1 Distill Llama 8B ist ein quelloffenes Sprachmodell von DeepSeek mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
If you are using the weights from this repository, please update to
Cosmos 1.0 Diffusion 7B Text2World ist ein quelloffenes Sprachmodell von NVIDIA mit 7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
More information please refer to our repo: https://github.com/VectorSpaceLab/OmniGen
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning
The CogAgent-9B-20241220 model is based on GLM-4V-9B, a bilingual open-source VLM base model. Through data collection and optimization, multi-stage training, and strategy improvements, CogAgent-9B-20241220 achieves significant advancements in GUI perception, inference prediction accuracy, action space completeness, and task generalizability. The model supports bilingual (Chinese and English) interaction with both screenshots and language input.
Granite Guardian 3.1 2B is a fine-tuned Granite 3.1 2B Instruct model designed to detect risks in prompts and responses. It can help with risk detection along many key dimensions catalogued in the IBM AI Risk Atlas. It is trained on unique data comprising human annotations and synthetic data informed by internal red-teaming. It outperforms other open-source models in the same space on standard benchmarks.
Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.
Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.
Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.
Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.
Using the 🤗's Diffusers library to run NOVA in a simple and efficient manner.
Our training data is an extension of the data used for Phi-3 and includes a wide variety of sources from:
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
[!WARNING] WARNING: This is a base language model that has not undergone instruction tuning or alignment with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications.
Model Summary: Granite-3.1-1B-A400M-Instruct is a 1B parameter long-context instruct model finetuned from Granite-3.1-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-3B-A800M-Instruct is a 3B parameter long-context instruct model finetuned from Granite-3.1-3B-A800M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-2B-Instruct is a 2B parameter long-context instruct model finetuned from Granite-3.1-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Aquila 135M Instruct ist ein quelloffenes Sprachmodell von BAAI mit 0.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Model Summary: Granite-Embedding-278M-Multilingual is a 278M parameter model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 768 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. This model is developed using contrastive finetuning, knowledge distillation and model merging for improved performance.
Model Summary: Granite-Embedding-107M-Multilingual is a 107M parameter dense biencoder embedding model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. This model is developed using contrastive finetuning, knowledge distillation and model merging for improved performance.
Model Summary: Granite-Embedding-30m-English is a 30M parameter dense bi-encoder embedding model from the Granite Embeddings suite that can be used to generate high quality text embeddings. This model produces embedding vectors of size 384 and is trained using a combination of open source relevance-pair datasets with permissive, enterprise-friendly license, and IBM collected and generated datasets. While maintaining competitive scores on academic benchmarks such as BEIR, this model also performs well on many enterprise use cases. This model is developed using retrieval oriented pre-training, c
News: Granite Embedding R2 models with 8192 context length released.
This repository provides the Depth ControlNet for Stable Diffusion 3.5 Large..
This repository provides the Blur ControlNet for Stable Diffusion 3.5 Large..
This repository provides the Canny ControlNet for Stable Diffusion 3.5 Large..
Install the transformers library from the source code:
Install the transformers library from the source code:
EuroLLM 9B Instruct ist ein quelloffenes Sprachmodell von Utter-project mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
EuroLLM 9B ist ein quelloffenes Sprachmodell von Utter-project mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma2 3b pt 896 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma2 3b pt 448 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma2 3b pt 224 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma2 3b ft docci 448 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma2 3b mix 224 ist ein multimodales Sprachmodell von Google mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
FLUX.1 Depth dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
FLUX.1 Canny dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Install the transformers library from the source code:
Install the transformers library from the source code:
Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 7B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.
Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 2B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
CogVideoX is an open-source video generation model similar to QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.
This model is the gptq-quantized version of Salamandra-2b for speculative decoding.
This model is the gptq-quantized version of Salamandra-7b for speculative decoding.
stable diffusion 3.5 medium ist ein multimodales Sprachmodell von Stabilityai mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
Below is the model card of Emu3-Chat model, which is adapted from the original Emu3 model card that you can find here.
Below is the model card of Emu3-Chat model, which is adapted from the original Emu3 model card that you can find here.
If you are using the weights from this repository, please update to
If you are using the weights from this repository, please update to
stable diffusion 3.5 large turbo ist ein multimodales Sprachmodell von Stabilityai mit 8.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
stable diffusion 3.5 large ist ein multimodales Sprachmodell von Stabilityai mit 8.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
This model hub includes a finetuned version of YOLOv8 and a finetuned BLIP-2 model on the above dataset respectively. For more details of the models used and finetuning, please refer to the paper.
This model is the DiT version of CogView3, a text-to-image generation model, supporting image generation from 512 to 2048px.
Model Summary: Granite-3.0-1B-A400M-Instruct is an 1B parameter model finetuned from Granite-3.0-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.0-1B-A400M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 8 trillion tokens sourced from diverse domains. During the second stage, it is further trained on 2 trillion tokens using a carefully curated mix of high-quality data, aiming to enhance its performance on specific tasks.
Model Summary: Granite-3.0-8B-Instruct is a 8B parameter model finetuned from Granite-3.0-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.0-8B-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 10 trillion tokens sourced from diverse domains. During the second stage, it is further trained on 2 trillion tokens using a carefully curated mix of high-quality data, aiming to enhance its performance on specific tasks.
Model Summary: Granite-3.0-2B-Instruct is a 2B parameter model finetuned from Granite-3.0-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Whisper is a state-of-the-art model for automatic speech recognition (ASR) and speech translation, proposed in the paper Robust Speech Recognition via Large-Scale Weak Supervision by Alec Radford et al. from OpenAI. Trained on 5M hours of labeled data, Whisper demonstrates a strong ability to generalise to many datasets and domains in a zero-shot setting.
This model is ready for non-commercial use.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
gemma 2 2b jpn it ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
--- license: apache-2.0 datasets: - projecte-aina/RAGMultilingual language: - es - en - ca libraryname: transformers ---
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---
Llama Guard 3 1B ist ein quelloffenes Sprachmodell von Meta mit 1.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 3B ist ein quelloffenes Sprachmodell von Meta mit 3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 3B Instruct ist ein quelloffenes Sprachmodell von Meta mit 3.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 1B Instruct ist ein quelloffenes Sprachmodell von Meta mit 1.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
OPI Llama 3.1 8B Instruct ist ein quelloffenes Sprachmodell von BAAI mit 8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
2b Version of model Sxxxxxx, without last epoch, instructed with baseline dataset including RAGMultilingual
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.
Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures.
CogVideoX is an open-source version of the video generation model originating from QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.
Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 model family, and the multimodal version comes with 128K context length (in tokens) it can support. The model underwent a rigorous enhancement process, incorporating both supervised fine-tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures.
🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]
LongWriter-glm4-9b is trained based on glm-4-9b, and is capable of generating 10,000+ words at once.
This instructed model uses a chat template that must be adhered to the input for conversational use. The easiest way to apply it is using the tokenizer's built-in chat template, as shown in the following snippet.
This is the model card for the first instruction tuned model of the EuroLLM series: EuroLLM-1.7B-Instruct. You can also check the pre-trained version: EuroLLM-1.7B.
This is the model card for the first pre-trained model of the EuroLLM series: EuroLLM-1.7B. You can also check the instruction tuned version: EuroLLM-1.7B-Instruct.
CogVideoX is an open-source version of the video generation model originating from QingYing. The table below displays the list of video generation models we currently offer, along with their foundational information.
FLUX.1 dev ist ein multimodales Sprachmodell von Black-forest-labs mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
ar stablelm 2 base ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
For more details please refer to our Github: FlagEmbedding.
Infinity-Instruct-7M-Gen-Llama3.1-70B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-7M and Infinity-Instruct-Gen and showing favorable results on AlpacaEval 2.0 and arena-hard compared to GPT4.
experimental7b rag instruct ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern und einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama Guard 3 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
experimental7b rag ist ein quelloffenes Sprachmodell von BSC-LT mit 7.8B Parametern und einem Kontextfenster von 8K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 8B Instruct ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
shieldgemma 9b ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
shieldgemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 405B Instruct ist ein quelloffenes Sprachmodell von Meta mit 405B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 70B Instruct ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2 2b it ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2 2b ist ein quelloffenes Sprachmodell von Google mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 405B ist ein quelloffenes Sprachmodell von Meta mit 405B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 70B ist ein quelloffenes Sprachmodell von Meta mit 70B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.1 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Infinity-Instruct-3M-0625-Llama3-8B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.
Infinity-Instruct-3M-0625-Yi-1.5-9B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.
Infinity-Instruct-3M-0625-Qwen2-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 and MT-Bench.
Infinity-Instruct-3M-0625-Mistral-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0625 and showing favorable results on AlpacaEval 2.0 compared to Mixtral 8x7B v0.1, Gemini Pro, and GPT-3.5.
We introduce CodeGeeX4-ALL-9B, the open-source version of the latest CodeGeeX4 model series. It is a multilingual code generation model continually trained on the GLM-4-9B, significantly enhancing its code generation capabilities. Using a single CodeGeeX4-ALL-9B model, it can support comprehensive functions such as code completion and generation, code interpreter, web search, function call, repository-level code Q&A, covering various scenarios of software development. CodeGeeX4-ALL-9B has achieved highly competitive performance on public benchmarks, such as BigCodeBench and NaturalCodeBench. I
Infinity-Instruct-3M-0613-Llama3-70B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0613 and showing favorable results on AlpacaEval 2.0 compared to GPT4-0613.
gemma 2 9b ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2 9b it ist ein quelloffenes Sprachmodell von Google mit 9.2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2 27b ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2 27b it ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Infinity-Instruct-3M-0613-Mistral-7B is an opensource supervised instruction tuning model without reinforcement learning from human feedback (RLHF). This model is just finetuned on Infinity-Instruct-3M and Infinity-Instruct-0613 and showing favorable results on AlpacaEval 2.0 compared to Mixtral 8x7B v0.1, Gemini Pro, and GPT-3.5.
In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.
In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.
stable diffusion 3 medium diffusers ist ein multimodales Sprachmodell von Stabilityai mit 2.1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
2024/08/12, 本仓库代码已更新并使用 transformers=4.44.0, 请及时更新依赖。
Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-2 showcased a nearly state-of-the-art performance among models with less than 13 billion parameters.
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the instruction-tuned 1.5B Qwen2 model.
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language model.
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:
Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
paligemma 3b ft cococap 448 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma 3b mix 224 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
paligemma 3b pt 224 ist ein multimodales Sprachmodell von Google mit 2.9B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
japanese stablelm 2 instruct 1 6b ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
japanese stablelm 2 base 1 6b ist ein quelloffenes Sprachmodell von Stabilityai mit 1.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
tokenizer = AutoTokenizer.frompretrained('nvidia/dragon-multiturn-query-encoder') queryencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-query-encoder') contextencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-context-encoder')
tokenizer = AutoTokenizer.frompretrained('nvidia/dragon-multiturn-query-encoder') queryencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-query-encoder') contextencoder = AutoModel.frompretrained('nvidia/dragon-multiturn-context-encoder')
Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.
Meta Llama Guard 2 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Meta Llama 3 8B ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Meta Llama 3 8B Instruct ist ein quelloffenes Sprachmodell von Meta mit 8B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Meta Llama 3 70B Instruct ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Meta Llama 3 70B ist ein quelloffenes Sprachmodell von Meta mit 71B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Stable LM 2 Chat 1.6B is a 1.6 billion parameter instruction tuned language model inspired by HugginFaceH4's Zephyr 7B training pipeline. The model is trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).
Stable LM 2 12B Chat is a 12 billion parameter instruction tuned language model trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).
Poro 34b chat is a chat-tuned version of Poro 34B trained to follow instructions in both Finnish and English. Quantized versions are available on Poro 34B-chat-GGUF.
gemma 1.1 2b it ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 1.1 7b it ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Stable LM 2 12B is a 12.1 billion parameter decoder-only language model pre-trained on 2 trillion tokens of diverse multilingual and code datasets for two epochs.
codegemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
This repository stores a development version of Stable LM 2 for sanity-checking/debugging the transformers implementation.
CodeLlama 34b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 34B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 70b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 13b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 13b Python hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 13b hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 7b Instruct hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
CodeLlama 7b hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
stable-code-instruct-3b is a 2.7B billion parameter decoder-only language model tuned from stable-code-3b. This model was trained on a mix of publicly available datasets, synthetic datasets using Direct Preference Optimization (DPO).
Viking 33B is a 33B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1300B billion as of this release). Viking 33B is a fully open source model and is made available under the Apache 2.0 License.
Viking 13B is a 13B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1.3 trillion as of this release). Viking 13B is a fully open source model and is made available under the Apache 2.0 License.
Viking 7B is a 7B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It has been trained on 2 trillion tokens. Viking 7B is a fully open source model and is made available under the Apache 2.0 License.
gemma 7b it ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 7b ist ein quelloffenes Sprachmodell von Google mit 8.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2b it ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 2b ist ein quelloffenes Sprachmodell von Google mit 2.5B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Bunny is a family of lightweight but powerful multimodal models. It offers multiple plug-and-play vision encoders, like EVA-CLIP, SigLIP and language backbones, including Phi-1.5, StableLM-2, Qwen1.5 and Phi-2. To compensate for the decrease in model size, we construct more informative training data by curated selection from a broader data source. Remarkably, our Bunny-3B model built upon SigLIP and Phi-2 outperforms the state-of-the-art MLLMs, not only in comparison with models of similar size but also against larger MLLM frameworks (7B), and even achieves performance on par with 13B models.
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
For more details please refer to our github repo: https://github.com/FlagOpen/FlagEmbedding
deepseek coder 7b instruct v1.5 ist ein quelloffenes Sprachmodell von DeepSeek mit 7B Parametern und einem Kontextfenster von 4K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Stable LM 2 Zephyr 1.6B is a 1.6 billion parameter instruction tuned language model inspired by HugginFaceH4's Zephyr 7B training pipeline. The model is trained on a mix of publicly available datasets and synthetic datasets, utilizing Direct Preference Optimization (DPO).
Please note: For commercial use, please refer to https://stability.ai/license
python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
Please note: For commercial use, please refer to https://stability.ai/license.
modelname = "deepseek-ai/deepseek-moe-16b-base" tokenizer = AutoTokenizer.frompretrained(modelname) model = AutoModelForCausalLM.frompretrained(modelname, torchdtype=torch.bfloat16, devicemap="auto") model.generationconfig = GenerationConfig.frompretrained(modelname) model.generationconfig.padtokenid = model.generationconfig.eostokenid
Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-2 showcased a nearly state-of-the-art performance among models with less than 13 billion parameters.
LlamaGuard 7b ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Introducing DeepSeek LLM, an advanced language model comprising 67 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
Please note: For commercial use, please refer to https://stability.ai/license.
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B, Aquila2-34B and Aquila2-70B-Expr , as well as AquilaChat2, the chat models, namely AquilaChat2-7B, AquilaChat2-34B and AquilaChat2-70B-Expr, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B, Aquila2-34B and Aquila2-70B-Expr , as well as AquilaChat2, the chat models, namely AquilaChat2-7B, AquilaChat2-34B and AquilaChat2-70B-Expr, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
通义千问-72B(Qwen-72B)是阿里云研发的通义千问大模型系列的720亿参数规模的模型。Qwen-72B是基于Transformer的大语言模型, 在超大规模的预训练数据上进行训练得到。预训练数据类型多样,覆盖广泛,包括大量网络文本、专业书籍、代码等。同时,在Qwen-72B的基础上,我们使用对齐机制打造了基于大语言模型的AI助手Qwen-72B-Chat。本仓库为Qwen-72B的仓库。
Please note: For commercial use, please refer to https://stability.ai/license.
codellama13b instruct 260k synthesis ist ein quelloffenes Sprachmodell von Stabilityai mit 13B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
japanese stable diffusion xl ist ein multimodales Sprachmodell von Stabilityai mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
A cute robot wearing a kimono writes calligraphy with one single brush — Stable Diffusion XL
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---
--- inference: false language: - en tags: - instruction-finetuning prettyname: JudgeLM-100K taskcategories: - text-generation ---
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Poro is a 34B parameter decoder-only transformer pretrained on Finnish, English and code. It was trained on 1 trillion tokens. Poro is a fully open source model and is made available under the Apache 2.0 License.
This is a 7B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese Stable LM Base Gamma 7B.
This is a 7B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, Mistral-7B-v0.1, to transfer the model's knowledge and capabilities to Japanese.
This is a 3B-parameter decoder-only Japanese language model fine-tuned on instruction-following datasets, built on top of the base model Japanese StableLM-3B-4E1T Base.
This is a 3B-parameter decoder-only language model with a focus on maximizing Japanese language modeling performance and Japanese downstream task performance. We conducted continued pretraining using Japanese data on the English language model, StableLM-3B-4E1T, to transfer the model's knowledge and capabilities to Japanese.
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
We opensource our Aquila2 series, now including Aquila2, the base language models, namely Aquila2-7B and Aquila2-34B, as well as AquilaChat2, the chat models, namely AquilaChat2-7B and AquilaChat2-34B, as well as the long-text chat models, namely AquilaChat2-7B-16k and AquilaChat2-34B-16k
More details please refer to our Github: FlagEmbedding.
StableLM-3B-4E1T is a 3 billion parameter decoder-only language model pre-trained on 1 trillion tokens of diverse English and code datasets for 4 epochs.
py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest
The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters. Mistral-7B-v0.1 outperforms Llama 2 13B on all benchmarks we tested.
We have updated the new reranker, supporting larger lengths, more languages, and achieving better performance.
More details please refer to our Github: FlagEmbedding.
For more details please refer to our Github: FlagEmbedding.
More details please refer to our Github: FlagEmbedding.
More details please refer to our Github: FlagEmbedding.
For more details please refer to our Github: FlagEmbedding.
For more details please refer to our Github: FlagEmbedding.
The language model Phi-1 is a Transformer with 1.3 billion parameters, specialized for basic Python coding. Its training involved a variety of data sources, including subsets of Python codes from The Stack v1.2, Q&A content from StackOverflow, competition code from codecontests, and synthetic Python textbooks and exercises generated by gpt-3.5-turbo-0301. Even though the model and the datasets are relatively small compared to contemporary Large Language Models (LLMs), Phi-1 has demonstrated an impressive accuracy rate exceeding 50% on the simple Python coding benchmark, HumanEval.
The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1, augmented with a new data source that consists of various NLP synthetic texts. When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-1.5 demonstrates a nearly state-of-the-art performance among models with less than 10 billion parameters.
StableCode-Completion-Alpha-3B-4K is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that topped the stackoverflow developer survey.
stablecode instruct alpha 3b ist ein quelloffenes Sprachmodell von Stabilityai mit 3.3B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Recommend switching to newest BAAI/bge-small-en-v1.5, which has more reasonable similarity distribution and same method of usage.
Recommend switching to newest BAAI/bge-base-en-v1.5, which has more reasonable similarity distribution and same method of usage.
Recommend switching to newest BAAI/bge-small-zh-v1.5, which has more reasonable similarity distribution and same method of usage.
Recommend switching to newest BAAI/bge-base-zh-v1.5, which has more reasonable similarity distribution and same method of usage.
Recommend switching to newest BAAI/bge-large-zh-v1.5, which has more reasonable similarity distribution and same method of usage.
Recommend switching to newest BAAI/bge-large-en-v1.5, which has more reasonable similarity distribution and same method of usage.
StableCode-Completion-Alpha-3B is a 3 billion parameter decoder-only code completion model pre-trained on diverse set of programming languages that were the top used languages based on the 2023 stackoverflow developer survey.
Use Stable Chat (Research Preview) to test Stability AI's best language models for free
Use Stable Chat (Research Preview) to test Stability AI's best language models for free
BF16/FP16版本|BF16/FP16 version codegeex2-6b
Use Stable Chat (Research Preview) to test Stability AI's best language models for free
Stable Beluga 1 is a Llama65B model fine-tuned on an Orca style Dataset
Aquila Language Model is the first open source language model that supports both Chinese and English knowledge, commercial license agreements, and compliance with domestic data regulations.
Llama 2 70b chat hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 2 7b chat hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 2 7b hf ist ein quelloffenes Sprachmodell von Meta mit 6.7B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 2 13b hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 2 13b chat hf ist ein quelloffenes Sprachmodell von Meta mit 13B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 2 70b hf ist ein quelloffenes Sprachmodell von Meta mit 69B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
OPI Galactica 6.7B ist ein quelloffenes Sprachmodell von BAAI mit 6.7B Parametern und einem Kontextfenster von 2K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
codeexecutor ist ein quelloffenes Sprachmodell von Microsoft mit 0.1B Parametern und einem Kontextfenster von 1K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
stable diffusion xl base 0.9 ist ein multimodales Sprachmodell von Stabilityai mit 2.6B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
StableLM-Tuned-Alpha is a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.
StableLM-Tuned-Alpha is a suite of 3B and 7B parameter decoder-only language models built on top of the StableLM-Base-Alpha models and further fine-tuned on various chat and instruction-following datasets.
📢 DISCLAIMER: The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.
📢 DISCLAIMER: The StableLM-Base-Alpha models have been superseded. Find the latest versions in the Stable LM Collection here.
ChatGLM-6B-INT4-QE 是 ChatGLM-6B 量化后的模型权重。具体的,ChatGLM-6B-INT4-QE 对 ChatGLM-6B 中的 28 个 GLM Block 、 Embedding 和 LM Head 进行了 INT4 量化。量化后的模型权重文件仅为 3G ,理论上 6G 显存(使用 CPU 即 6G 内存)即可推理,具有在嵌入式设备(如树莓派)上运行的可能。
This model was previously named "PubMedELECTRA large (abstracts)". You can either adopt the new model name "microsoft/BiomedNLP-BiomedELECTRA-large-uncased-abstract" or update your transformers library to version 4.22+ if you need to refer to the old name.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
Whisper is a pre-trained model for automatic speech recognition (ASR) and speech translation. Trained on 680k hours of labelled data, Whisper models demonstrate a strong ability to generalise to many datasets and domains without the need for fine-tuning.
This is the retrieval model for ReACC: A Retrieval-Augmented Code Completion Framework.
- Developed by: Microsoft Team - Shared by [Optional]: Hugging Face - Model type: Feature Engineering - Language(s) (NLP): en - License: Apache-2.0 - Related Models: - Parent Model: RoBERTa - Resources for more information: - Associated Paper
- Developed by: Microsoft Team - Shared by [Optional]: Hugging Face - Model type: Feature Engineering - Language(s) (NLP): en - License: Apache-2.0 - Related Models: - Parent Model: RoBERTa - Resources for more information: - Associated Paper
unixcoder base unimodal ist ein quelloffenes Sprachmodell von Microsoft mit einem Kontextfenster von 1K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
This model uses a BERT large architecture [1] pretrained from scratch using the Wikipedia [2], Common Crawl [3], PMINDIA [4] and Dakshina [5] corpora for 17 [6] Indian languages.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
CodeGPT small py ist ein quelloffenes Sprachmodell von Microsoft, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
codebert base ist ein quelloffenes Sprachmodell von Microsoft mit einem Kontextfenster von 514 Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 30B A3B Instruct 2507 ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 4B Thinking 2507 ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek R1 Distill 1.5B ist ein quelloffenes Sprachmodell von DeepSeek mit 1.5B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 Coder Next ist ein quelloffenes Sprachmodell von Qwen mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Open-weight video generation (text to video). We host it for you on dedicated GPUs in an EU datacenter. Contact us for a quote.
translategemma 4b ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Super Loes (Kimi K3) ist ein quelloffenes Sprachmodell von Hyai mit einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Open-weight image generation (text to image). We host it for you on a dedicated GPU in an EU datacenter. Contact us for a quote.
Qwen3Guard Gen 0.6B ist ein quelloffenes Sprachmodell von Qwen mit 0.6B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 30B A3B ist ein quelloffenes Sprachmodell von Qwen mit 30B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 8B Thinking ist ein quelloffenes Sprachmodell von Qwen mit 8B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen3 VL 2B ist ein quelloffenes Sprachmodell von Qwen mit 2B Parametern und einem Kontextfenster von 262K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 2.5 Coder 1.5B ist ein quelloffenes Sprachmodell von Qwen mit 1.5B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen2.5 VL 3B ist ein quelloffenes Sprachmodell von Qwen mit 3B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
deepseek coder 6.7b ist ein quelloffenes Sprachmodell von DeepSeek mit 6.7B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3n E2B ist ein quelloffenes Sprachmodell von Google mit 2B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Gemma 3 4B ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Gemma 3 1B ist ein quelloffenes Sprachmodell von Google mit 1B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
gemma 3 12b ist ein quelloffenes Sprachmodell von Google mit 12B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
functiongemma 270m ist ein quelloffenes Sprachmodell von Google, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Open-weight image generation (text to image), served on a dedicated GPU in an EU datacenter via an OpenAI-compatible /v1/images/generations endpoint. Deploy it as a dedicated instance from the wizard, or contact us for help sizing it.
DeepSeek Coder V2 Lite (16B) ist ein quelloffenes Sprachmodell von DeepSeek mit 16B Parametern und einem Kontextfenster von 164K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek V3 (685B MoE) ist ein quelloffenes Sprachmodell von DeepSeek mit 685B Parametern und einem Kontextfenster von 164K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Open-weight text-to-speech. We host it for you on a dedicated GPU in an EU datacenter, served via an OpenAI-compatible /v1/audio/speech endpoint. Contact us for a quote.
Qwen2.5 VL 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 128K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 2.5 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen2.5 Coder 14B ist ein quelloffenes Sprachmodell von Qwen mit 14B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen2.5 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 3 4B ist ein quelloffenes Sprachmodell von Qwen mit 4B Parametern und einem Kontextfenster von 41K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 2.5 Coder 7B ist ein quelloffenes Sprachmodell von Qwen mit 7B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 2.5 Coder 32B ist ein quelloffenes Sprachmodell von Qwen mit 32B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
DeepSeek R1 Distill 32B ist ein quelloffenes Sprachmodell von DeepSeek mit 32B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Qwen 2.5 72B ist ein quelloffenes Sprachmodell von Qwen mit 72B Parametern und einem Kontextfenster von 33K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Phi 4 Mini (3.8B) ist ein quelloffenes Sprachmodell von Microsoft mit 3.8B Parametern und einem Kontextfenster von 131K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 11B Vision ist ein quelloffenes Sprachmodell von Meta mit 11B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Phi 4 (14B) ist ein quelloffenes Sprachmodell von Microsoft mit 14B Parametern und einem Kontextfenster von 16K Tokens, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
medgemma 4b ist ein quelloffenes Sprachmodell von Google mit 4B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
medgemma 27b ist ein quelloffenes Sprachmodell von Google mit 27B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Sovereign EU model fine-tuned by HostYourAI on dutch-clean.
Llama 4 Maverick (17Bx128E) ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 3.2 90B Vision ist ein quelloffenes Sprachmodell von Meta mit 90B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Llama 4 Scout (17Bx16E) ist ein quelloffenes Sprachmodell von Meta mit 17B Parametern, gehostet auf europäischen GPUs über eine OpenAI-kompatible API.
Keine Kreditkarte nötig. Pay as you go, jederzeit kündbar.
Heute kostenlos starten