551 open-source modellen, gehost op GPU's in de EU. Eén OpenAI-compatibele API-key, scale-to-zero of dedicated.
551 modellen
This repository contains an FP8 quantized version of the Qwen3-VL-8B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!
This repository contains an FP8 quantized version of the Qwen3-VL-32B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!
Qwen3 VL 8B is een open-source taalmodel van Qwen met 8B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
This repository contains an FP8 quantized version of the Qwen3-VL-30B-A3B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Qwen3 VL 2B is een open-source taalmodel van Qwen met 2B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Loes Large World: Qwen3.5-27B dense base met volledig-Europese SFT-adapter (LoRA), geserveerd via vLLM --enable-lora + qwen3 reasoning-parser, CUDA-graphs. Topmodel-spoor voor chat.loes.ai.
Note: DeepSeek-V4-Flash-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec
Qwen3.6 27B is een open-source taalmodel van Qwen met 27B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 VL 4B is een open-source taalmodel van Qwen met 4B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek R1 Distill 32B is een open-source taalmodel van DeepSeek met 32B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 VL 32B is een open-source taalmodel van Qwen met 32B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek R1 Distill Llama 70B is een open-source taalmodel van DeepSeek met 70B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Holo2 30B A3B is een open-source taalmodel van Holo2-30b-a3b met een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 30B A3B is een open-source taalmodel van Qwen met 30B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 235B A22B Instruct is een open-source taalmodel van Qwen3-235b-a22b-instruct-2507 met een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 3 Coder 30B-A3B (MoE) is een open-source taalmodel van Qwen met 30B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
GPT-OSS 120B is een open-source taalmodel van Gpt-oss-120b met een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
EGM-Qwen3-VL-8B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-8B-Thinking.
Gemma 4 26B A4B is een open-source taalmodel van Gemma-4-26b-a4b-it met een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
We introduce X-Reasoner, a vision-language model posttrained solely on general-domain text for generalizable reasoning, using a twostage approach: an initial supervised fine-tuning phase with distilled long chainof-thoughts, followed by reinforcement learning with verifiable rewards. Experiments show that X-Reasoner successfully transfers reasoning capabilities to both multimodal and out-of-domain settings, outperforming existing state-of-theart models trained with in-domain and multimodal data across various general and medical benchmarks. More details can be found in the paper: X-Reasoner: T
gemma 3 27b it is een multimodaal taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Devstral 2 123B is een open-source taalmodel van Devstral-2-123b-instruct-2512 met een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Phi-3.5-vision is a lightweight, state-of-the-art open multimodal model built upon datasets which include - synthetic data and filtered publicly available websites - with a focus on very high-quality, reasoning dense data both on text and vision. The model belongs to the Phi-3 model family, and the multimodal version comes with 128K context length (in tokens) it can support. The model underwent a rigorous enhancement process, incorporating both supervised fine-tuning and direct preference optimization to ensure precise instruction adherence and robust safety measures.
Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.
Model Summary: Granite Vision 4.1 4B is a vision-language model (VLM) that delivers frontier-level performance on structured document extraction tasks — chart extraction, table extraction, and semantic key-value pair extraction — in a compact 4B parameter footprint, providing a lightweight alternative to much larger frontier models for these tasks:
Salamandra-VL-7B-2512 is the latest version of the Salamandra vision model family. This version brings significant improvements in both architecture and training data.
Model Summary: granite-vision-3.1-2b-preview is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model (https://huggingface.co/ibm-granite/granite-3.1-2b-instruct) with
Qwen3.5 397B A17B is een open-source taalmodel van Qwen3.5-397b-a17b met een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Gemma 3 27B is een open-source taalmodel van Gemma-3-27b-it met een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Inference using Huggingface transformers on NVIDIA GPUs. Requirements tested on python 3.12.9 + CUDA11.8:
py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest
DeepSeek R1 Distill 14B is een open-source taalmodel van DeepSeek met 14B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 8B is een open-source taalmodel van Qwen met 8B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 4 31B is een open-source taalmodel van Google met 31B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3.5 27B is een open-source taalmodel van Qwen met 27B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
GLM-4.7-Flash is a 30B-A3B MoE model. As the strongest model in the 30B class, GLM-4.7-Flash offers a new option for lightweight deployment that balances performance and efficiency.
Llama 3.1 8B Instruct is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 Coder 30B A3B is een open-source taalmodel van Qwen3-coder-30b-a3b-instruct met een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.3 70B Instruct is een open-source taalmodel van Llama-3.3-70b-instruct met een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
Gemma is a family of open models built by Google DeepMind. Gemma 4 models are multimodal, handling text and image input (with audio supported on E2B, E4B, and 12B) and generating text output. This release includes open-weights models in both pre-trained and instruction-tuned variants. Gemma 4 features a context window of up to 256K tokens and maintains multilingual support in over 140 languages.
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
Qwen3.6 27B FP8 is een open-source taalmodel van Qwen met 27B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Mistral Small 3.2 24B is een open-source taalmodel van Mistral-small-3.2-24b-instruct-2506 met een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3.5 35B A3B is een open-source taalmodel van Qwen met 35B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
GLM 5.2 (always-on) is een open-source taalmodel van Glm-5.2 met een contextvenster van 203K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Pixtral 12B is een open-source taalmodel van Pixtral-12b-2409 met een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek R1 0528 Qwen3 8B is een open-source taalmodel van DeepSeek met 8B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Voxtral Small 24B is een open-source taalmodel van Voxtral-small-24b-2507 met een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Kimi K3 is Moonshot AI's new flagship open-weight model: a 2.8 trillion parameter Mixture-of-Experts with a 1 million token context window, built for coding and long-horizon agent work. Moonshot has announced the open weights for late July 2026, with vLLM support from day one. Talk to us about running it on European infrastructure.
This model is ready for commercial or non-commercial use. <br
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
GELab Zero 4B preview Sico Evolution is een multimodaal taalmodel van Microsoft met 4.4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.
Note: DeepSeek-V4-Pro-DSpark is not a new model. It is the same checkpoint with an additional speculative decoding module attached. A minimal inference example is available in the inference folder. For more details, refer to: https://github.com/deepseek-ai/DeepSpec
This model is ready for commercial or non-commercial use. <br
This model is ready for commercial or non-commercial use. <br
We're introducing GLM-5.2, our latest flagship model for long-horizon tasks. It marks a substantial leap in long-horizon task capability over its predecessor GLM-5.1 and, for the first time, delivers that capability on a solid 1M-token context. GLM-5.2's new capabilities include: - Solid 1M Context: A solid 1M-token context that stably sustains long-horizon work - Advanced Coding with Flexible Effort: Stronger coding capabilities with multiple thinking effort levels to balance performance and latency - Improved Architecture: We propose IndexShare, which reuses the same indexer across every fou
Kimi K2.7 Code is a coding-focused agentic model built upon Kimi K2.6. With substantial improvements on real-world long-horizon coding tasks, it strengthens end-to-end task completion across complex software engineering workflows while improving token efficiency, reducing thinking-token usage by approximately 30% compared with Kimi K2.6.
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
For more details on how to deploy and use the model - see the Quick Start Guide below!
For more details on how to deploy and use the model - see the Quick Start Guide below!
This model is ready for commercial/non-commercial use. <br
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
This model is ready for commercial/non-commercial use. <br
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress: New versions will be released over the coming months.
This model is ready for commercial/non-commercial use. <br
This model is ready for commercial/non-commercial use. <br
[!NOTE] WARNING: This model has been trained on instructions but has not undergone safety or value alignment. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
This model is ready for commercial/non-commercial use. <br
NVIDIA Cosmos Reason 2 is an open, customizable, 32B-parameter reasoning vision language model (VLM) for physical AI and robotics that enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
[!Note] This model card is for the new versions of the Gemma 4 family optimized with Quantization-Aware Training (QAT), which allows preserving similar quality to bfloat16 while dramatically reducing the memory requirements to load the model. Four versions of the QAT checkpoints are available: Unquantized QAT checkpoints (Q40): Half-precision weights extracted from the QAT pipeline, ideal for custom downstream compilation and research. Available for Gemma 4 E2B, E4B, 12B, 26B A4B, and 31B, and their drafter models. GGUF (Q40): Ready-to-deploy formats for broad ecosystem compatibility. Availabl
[!NOTE] This repository contains the FP8 version of 4.1-30B. Please refer to the the original instruct model's model card for additional details: https://huggingface.co/ibm-granite/granite-4.1-30b
[!NOTE] This repository contains the FP8 version of 4.1-8B. Please refer to the the original instruct model's model card for additional details: https://huggingface.co/ibm-granite/granite-4.1-8b
Granite Guardian 4.1 8B introduces improved Bring Your Own Criteria (BYOC) support, enabling users to define arbitrary judging criteria beyond the pre-baked safety and hallucination detectors. The model can now faithfully evaluate complex, multi-part requirements such as formatting rules, length constraints, and domain-specific instructions.
Qwen3.6 35B A3B is een open-source taalmodel van Qwen met 35B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Kimi K2.6 is an open-source, native multimodal agentic model that advances practical capabilities in long-horizon coding, coding-driven design, proactive autonomous execution, and swarm-based task orchestration.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
Model Summary: Granite‑4.1‑30B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of text‑to‑text generation tasks. In addition to standard generation, it supports Fill‑in‑the‑Middle (FIM) code completion through specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens using a five‑phase training strategy: 10 trillion tokens in phase one, 2 trillion tokens each in phases two and three, and 0.5 trillion tokens in phase four. In the final phase, long‑context extension is applied to expand the
Model Summary: Granite-4.1-30B is a 30B parameter long-context instruct model finetuned from Granite-4.1-30B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
Model Summary: Granite‑4.1‑8B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of text‑to‑text generation tasks. In addition to standard generation, it supports Fill‑in‑the‑Middle (FIM) code completion through specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens using a five‑phase training strategy: 10 trillion tokens in phase one, 2 trillion tokens each in phases two and three, and 0.5 trillion tokens in phase four. In the final phase, long‑context extension is applied to expand the m
Model Summary: Granite-4.1-8B is a 8B parameter long-context instruct model finetuned from Granite-4.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
Model Summary: Granite‑4.1‑3B‑Base is a decoder‑only language model with long‑context capabilities, designed to support a broad range of general text‑to‑text generation tasks, as well as fill‑in‑the‑Middle (FIM) code completion. This model shares the same underlying architecture and weights as Granite 4.0 3B Micro, which is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the third, and 0.5 trillion in the final stage. An additional training phase is applied
Model Summary: Granite-4.1-3B is a 3B parameter long-context instruct model finetuned from Granite-4.1-3B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. Granite 4.1 models have gone through an improved post-training pipeline, including supervised finetuning and reinforcement learning alignment, resulting in enhanced tool calling, instruction following, and chat capabilities.
GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).
GLM-5.1 is our next-generation flagship model for agentic engineering, with significantly stronger coding capabilities than its predecessor. It achieves state-of-the-art performance on SWE-Bench Pro and leads GLM-5 by a wide margin on NL2Repo (repo generation) and Terminal-Bench 2.0 (real-world terminal tasks).
This model is ready for commercial/non-commercial use. <br
EGM-Qwen3-VL-4B-SFT is the supervised fine-tuning (SFT) checkpoint from the first stage of the EGM (Efficient Visual Grounding Language Models) training pipeline. It is built on top of Qwen3-VL-4B-Thinking.
EGM-Qwen3-VL-4B is an efficient visual grounding model from the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-4B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
This model is a fine-tuned derivative of Qwen3.5-35B-A3B. Follow the Qwen3.5-35B-A3B serving guide for deployment with vLLM, replacing the model path with nvidia/NVIDIA-Ising-Calibration-1-35B-A3B. Suggested inference settings: temperature=0.2, maxtokens=16384.
- Nemotron-Cascade-2-30B-A3B follows the ChatML template and supports both thinking and instruct (non-thinking) modes. Reasoning content is enclosed within <think and </think tags. To activate the instruct (non-thinking) mode, we prepend <think</think to the beginning of the assistant’s response.
Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.
Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.
Use temperature=1.0 and topp=0.95 across all tasks and serving backends — reasoning, tool calling, and general chat alike.
The pretraining data has a cutoff date of September 2024\.
Model Dates: Trained between Oct 2025 and March 2026
Model Summary: Granite-4.0-3B-Vision is a vision-language model (VLM) designed for enterprise-grade document data extraction. It focuses on specialized, complex extraction tasks that ultracompact models often struggle with:
EGM-Qwen3-VL-8B is the flagship model of the EGM (Efficient Visual Grounding Language Models) family. It is built on top of Qwen3-VL-8B-Thinking and trained with a two-stage pipeline: supervised fine-tuning (SFT) followed by reinforcement learning (RL) using GRPO (Group Relative Policy Optimization).
[!Note] This repository contains int4-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc.
[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger
Qwen3.5 0.8B is een open-source taalmodel van Qwen met 0.8B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3.5 2B is een open-source taalmodel van Qwen met 2B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3.5 4B is een open-source taalmodel van Qwen met 4B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3.5 9B is een open-source taalmodel van Qwen met 9B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
[!Note] This repository contains model weights and configuration files for the pre-trained only model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, etc. The intended use cases are fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction. However, the control tokens, e.g., <|imstart| and <|imend| were trained to allow efficient LoRA-style PEFT with the official chat template, mitigating the need to finetune embeddings, a significant optimization given Qwen3.5's larger
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
This model is ready for commercial/non-commercial use. <br
Qwen3.5 122B A10B is een open-source taalmodel van Qwen met 122B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
This model is ready for commercial use. <br
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
[!Note] This repository contains FP8-quantized model weights and configuration files for the post-trained model in the Hugging Face Transformers format. These artifacts are compatible with Hugging Face Transformers, vLLM, SGLang, KTransformers, etc. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original model.
This model is ready for commercial/non-commercial use. <br
We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.
We are launching GLM-5, targeting complex systems engineering and long-horizon agentic tasks. Scaling is still one of the most important ways to improve the intelligence efficiency of Artificial General Intelligence (AGI). Compared to GLM-4.5, GLM-5 scales from 355B parameters (32B active) to 744B parameters (40B active), and increases pre-training data from 23T to 28.5T tokens. GLM-5 also integrates DeepSeek Sparse Attention (DSA), largely reducing deployment cost while preserving long-context capacity.
Today, we're announcing Qwen3-Coder-Next-FP8, an open-weight language model designed specifically for coding agents and local development. It features the following key enhancements:
This model is ready for commercial/non-commercial use. <br
Qwen3 Coder Next is een open-source taalmodel van Qwen met een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
GLM-OCR is a multimodal OCR model for complex document understanding, built on the GLM-V encoder–decoder architecture. It introduces Multi-Token Prediction (MTP) loss and stable full-task reinforcement learning to improve training efficiency, recognition accuracy, and generalization. The model integrates the CogViT visual encoder pre-trained on large-scale image–text data, a lightweight cross-modal connector with efficient token downsampling, and a GLM-0.5B language decoder. Combined with a two-stage pipeline of layout analysis and parallel recognition based on PP-DocLayout-V3, GLM-OCR deliver
Chart2CSV is a specialized vision-language model fine-tuned for the accurate extraction of tabular data from charts and visualizations. Built on top of ibm-granite/granite-vision-3.3-2b, it produces machine-readable CSV outputs with improved numeric fidelity compared to general-purpose VLMs. The model is trained using code-guided synthetic chart data following the ChartGen methodology, which strengthens factual grounding and reduces hallucination in the Chart-to-CSV task.
[!NOTE] WARNING: Although this model has undergone safety and value alignment, it may still occasionally generate unintended or undesired outputs. Work In Progress New versions will be available during the coming weeks/months. Sampling Parameters: For optimal performance, we recommend using temperatures close to zero (0 - 0.2). Additionally, we advise against using any type of repetition penalty, as from our experience, it negatively impacts instructed model's responses.
This is the model card for EuroLLM-9B-Instruct-2512, an improved version of utter-project/EuroLLM-9B-Instruct. In comparison with the previous version, this version includes the long-context extension phase and the revamped post-training recipe from utter-project/EuroLLM-22B-Instruct.
translategemma 27b it is een multimodaal taalmodel van Google met 29B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
translategemma 12b it is een multimodaal taalmodel van Google met 13B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
translategemma 4b it is een multimodaal taalmodel van Google met 5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
medgemma 1.5 4b it is een multimodaal taalmodel van Google met 4.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Kimi K2.5 is an open-source, native multimodal agentic model built through continual pretraining on approximately 15 trillion mixed visual and text tokens atop Kimi-K2-Base. It seamlessly integrates vision and language understanding with advanced agentic capabilities, instant and thinking modes, as well as conversational and agentic paradigms.
GLM-4.7, your new coding partner, is coming with the following features:
GLM-4.7, your new coding partner, is coming with the following features:
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
This is the model card for EuroMoE-2.6B-A0.6B-Instruct-2512. You can also check the pre-trained version: EuroMoE-2.6B-A0.6B-2512.
Cosmos Reason2 2B is een multimodaal taalmodel van NVIDIA met 2.4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Cosmos Reason2 8B is een multimodaal taalmodel van NVIDIA met 8.8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
OptiMind-SFT is a specialized 20B parameter model designed to bridge the gap between natural language and executable optimization solvers. It automates the translation of complex decision-making problems—such as supply chain planning, scheduling, and resource allocation—into correct MILP formulations.
⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.
⚠️ This project is intended for research and educational purposes only. Any use for illegal data access, system interference, or unlawful activities is strictly prohibited. Please review our Terms of Use carefully.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
This is the model card for EuroLLM-22B-Instruct. You can also check the pre-trained version: EuroLLM-22B-2515.
The post-training data has a cutoff date of November 28, 2025\. The pre-training data has a cutoff date of June 25, 2025\.
We introduce DeepSeek-V3.2, a model that harmonizes high computational efficiency with superior reasoning and agent performance. Our approach is built upon three key technical breakthroughs:
This is the model card for EuroLLM-22B. You can also check the post-trained version: EuroLLM-22B-Instruct-2515.
NVIDIA Nemotron Parse v1.1 is designed to understand document semantics and extract text and tables elements with spatial grounding. Given an image, NVIDIA Nemotron Parse v1.1 produces structured annotations, including formatted text, bounding-boxes and the corresponding semantic classes, ordered according to the document's reading flow. It overcomes the shortcomings of traditional OCR technologies that struggle with complex document layouts with structural variability, and helps transform unstructured documents into actionable and machine-usable representations. This has several downstream be
- Repository: https://github.com/zheny2751-dotcom/WebVIA - Paper: https://arxiv.org/pdf/2511.06251
- Repository: https://github.com/zai-org/UI2CodeN - Paper: https://arxiv.org/abs/2511.08195
Kimi K2 Thinking is the latest, most capable version of open-source thinking model. Starting with Kimi K2, we built it as a thinking agent that reasons step-by-step while dynamically invoking tools. It sets a new state-of-the-art on Humanity's Last Exam (HLE), BrowseComp, and other benchmarks by dramatically scaling multi-step reasoning depth and maintaining stable tool-use across 200–300 sequential calls. At the same time, K2 Thinking is a native INT4 quantization model with 256k context window, achieving lossless reductions in inference latency and GPU memory usage.
Description: Fara-7B is Microsoft's first agentic small language model (SLM) designed specifically for computer use. With only 7 billion parameters, Fara-7B is an ultra-compact Computer Use Agent (CUA) that achieves state-of-the-art performance within its size class and is competitive with larger, more resource-intensive agentic systems.
Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.
Kimi Linear is a hybrid linear attention architecture that outperforms traditional full attention methods across various contexts, including short, long, and reinforcement learning (RL) scaling regimes. At its core is Kimi Delta Attention (KDA)—a refined version of Gated DeltaNet that introduces a more efficient gating mechanism to optimize the use of finite-state RNN memory.
This model is for research and development only. <br
- Repository: https://github.com/thu-coai/Glyph - Paper: https://arxiv.org/abs/2510.17800
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
torch==2.6.0 transformers==4.46.3 tokenizers==0.20.3 einops addict easydict pip install flash-attn==2.7.3 --no-build-isolation
[!NOTE] This repository contains the FP8 version of Granite-3.3-8b-Instruct. Please reference the base model's full model card here: https://huggingface.co/ibm-granite/granite-3.3-8b-instruct
This model is for research and development only.
This repository contains an FP8 quantized version of the Qwen3-VL-4B-Instruct model. The quantization method is fine-grained fp8 quantization with block size of 128, and its performance metrics are nearly identical to those of the original BF16 model. Enjoy!
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
functiongemma 270m it is een open-source taalmodel van Google met 0.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Model Summary: Granite-4.0-350M-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the
Model Summary: Granite-4.0-350M is a lightweight instruct model finetuned from Granite-4.0-350M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
Model Summary: Granite-4.0-1B-Base is a lightweight decoder-only language model designed for scenarios where efficiency and speed are critical. They can run on resource-constrained devices such as smartphones or IoT hardware, enabling offline and privacy-preserving applications. It also supports Fill-in-the-Middle (FIM) code completion through the use of specialized prefix and suffix tokens. The model is trained from scratch on approximately 15 trillion tokens following a four-stage training strategy: 10 trillion tokens in the first stage, 2 trillion in the second, another 2 trillion in the th
Model Summary: Granite-4.0-1B is a lightweight instruct model finetuned from Granite-4.0-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
Model Summary: Granite-4.0-H-350M is a lightweight instruct model finetuned from Granite-4.0-H-350M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
Model Summary: Granite-4.0-H-1B is a lightweight instruct model finetuned from Granite-4.0-H-1B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques including supervised finetuning, reinforcement learning, and model merging.
Unlike typical LLMs that are trained to play the role of the "assistant" in conversation, we trained UserLM-8b to simulate the “user” role in conversation (by training it to predict user turns in a large corpus of conversations called WildChat). This model is useful in simulating more realistic conversations, which is in turn useful in the development of more robust assistants.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
Compared with GLM-4.5, GLM-4.6 brings several key improvements:
[!WARNING] WARNING: This is a language model that has undergone instruction tuning for conversational settings that exploit function calling capabilities. It has not been aligned with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications. How to use from datetime import datetime from transformers import AutoTokenizer, AutoModelForCausalLM import transformer
Compared with GLM-4.5, GLM-4.6 brings several key improvements:
We are excited to announce the official release of DeepSeek-V3.2-Exp, an experimental version of our model. As an intermediate step toward our next-generation architecture, V3.2-Exp builds upon V3.1-Terminus by introducing DeepSeek Sparse Attention—a sparse attention mechanism designed to explore and validate optimizations for training and inference efficiency in long-context scenarios.
This update maintains the model's original capabilities while addressing issues reported by users, including:
The pretraining data has a cutoff date of September 2024.
Meet Qwen3-VL — the most powerful vision-language model in the Qwen series to date.
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.
gpt-oss-safeguard-120b and gpt-oss-safeguard-20b are safety reasoning models built-upon gpt-oss. With these models, you can classify text content based on safety policies that you provide and perform a suite of foundational safety tasks. These models are intended for safety use cases. For other applications, we recommend using gpt-oss models.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
📣 Update [10-07-2025]: Added a default system prompt to the chat template to guide the model towards more professional, accurate, and safe responses.
This model is ready for commercial/non-commercial use. <br
This model is ready for commercial/non-commercial use. <br
This model is ready for commercial/non-commercial use. <br
Kimi K2-Instruct-0905 is the latest, most capable version of Kimi K2. It is a state-of-the-art mixture-of-experts (MoE) language model, featuring 32 billion activated parameters and a total of 1 trillion parameters.
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:
TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.
DeepSeek-V3.1 is a hybrid model that supports both thinking mode and non-thinking mode. Compared to the previous version, this upgrade brings improvements in multiple aspects:
1. Model Summary 2. How to use 3. Evaluation 4. Training 5. Limitations 6. Legal Aspects
The pretraining data has a cutoff date of September 2024.
TowerVision is a family of open-source multilingual vision-language models with strong capabilities optimized for a variety of vision-language use cases, including image captioning, visual understanding, summarization, question answering, and more. TowerVision excels particularly in multimodal multilingual translation benchmarks and culturally-aware tasks, demonstrating exceptional performance across 20 languages and dialects.
This model is part of the GLM-V family of models, introduced in the paper GLM-4.1V-Thinking and GLM-4.5V: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.
Vision-language models (VLMs) have become a key cornerstone of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs urgently need to enhance reasoning capabilities beyond basic multimodal perception — improving accuracy, comprehensiveness, and intelligence — to enable complex problem solving, long-context understanding, and multimodal agents.
We introduce the updated version of the Qwen3-4B-FP8 non-thinking mode, named Qwen3-4B-Instruct-2507-FP8, featuring the following key enhancements:
gemma 3 270m is een open-source taalmodel van Google met 0.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 4B Instruct 2507 is een open-source taalmodel van Qwen met 4B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Welcome to the gpt-oss series, OpenAI’s open-weight models designed for powerful reasoning, agentic tasks, and versatile developer use cases.
Llama-3.3-Nemotron-Super-49B-v1.5-FP8 is a significantly upgraded version of Llama-3.3-Nemotron-Super-49B-v1 and is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and agentic tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct-FP8. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
Qwen3-Coder is available in multiple sizes. Today, we're excited to introduce Qwen3-Coder-30B-A3B-Instruct. This streamlined model maintains impressive performance and efficiency, featuring the following key enhancements:
gemma 3 270m it is een open-source taalmodel van Google met 0.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
Compared to nvidia/DeepSeek-R1-0528-FP4, this checkpoint additionally quantizes the wo module in attention layers.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
We present GLM-4.5, an open-source Mixture-of-Experts (MoE) large language model with 355B total parameters and 32B activated parameters, featuring a hybrid reasoning method that supports both thinking and direct response modes. Through multi-stage training on 23T tokens and comprehensive post-training with expert model iteration and reinforcement learning, GLM-4.5 achieves strong performance across agentic, reasoning, and coding (ARC) tasks, scoring 70.1% on TAU-Bench, 91.0% on AIME 24, and 64.2% on SWE-bench Verified. With much fewer parameters than several competitors, GLM-4.5 ranks 3rd ove
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The GLM-4.5 series models are foundation models designed for intelligent agents. GLM-4.5 has 355 billion total parameters with 32 billion active parameters, while GLM-4.5-Air adopts a more compact design with 106 billion total parameters and 12 billion active parameters. GLM-4.5 models unify reasoning, coding, and intelligent agent capabilities to meet the complex demands of intelligent agent applications.
The MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e. PubMed commercial, Medical Wikipedia, Medical Guidelines, Medical Coding, and open-source clinical documents) and merged back with the SLERP method in their base model to conserve general abilities. One model combined all five experts into one general expert with the multi-model merging method BreadCrumbs. F
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
medgemma 27b it is een multimodaal taalmodel van Google met 29B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Dayhoff is an Atlas of both protein sequence data and generative language models — a centralized resource that brings together 3.34 billion protein sequences across 1.7 billion clusters of metagenomic and natural protein sequences (GigaRef), 46 million structure-derived synthetic sequences (BackboneRef), and 16 million multiple sequence alignments (OpenProteinSet). These models can natively predict zero-shot mutation effects on fitness, scaffold structural motifs by conditioning on evolutionary or structural context, and perform guided generation of novel proteins within specified families. Le
Kimi K2 is a state-of-the-art mixture-of-experts (MoE) language model with 32 billion activated parameters and 1 trillion total parameters. Trained with the Muon optimizer, Kimi K2 achieves exceptional performance across frontier knowledge, reasoning, and coding tasks while being meticulously optimized for agentic capabilities.
Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.
Vision-Language Models (VLMs) have become foundational components of intelligent systems. As real-world AI tasks grow increasingly complex, VLMs must evolve beyond basic multimodal perception to enhance their reasoning capabilities in complex tasks. This involves improving accuracy, comprehensiveness, and intelligence, enabling applications such as complex problem solving, long-context understanding, and multimodal agents.
Phi-tiny-MoE is a lightweight Mixture of Experts (MoE) model with 3.8B total parameters and 1.1B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a larger variant, Phi-mini-MoE, with 7.6B total and 2.4B activated pa
Phi-mini-MoE is a lightweight Mixture of Experts (MoE) model with 7.6B total parameters and 2.4B activated parameters. It is compressed and distilled from the base model shared by Phi-3.5-MoE and GRIN-MoE using the SlimMoE approach, then post-trained via supervised fine-tuning and direct preference optimization for instruction following and safety. The model is trained on Phi-3 synthetic data and filtered public documents, with a focus on high-quality, reasoning-dense content. It is part of the SlimMoE series, which includes a smaller variant, Phi-tiny-MoE, with 3.8B total and 1.1B activated p
[!Note] This is an improved version of Kimi-VL-A3B-Thinking. Please consider using this updated model instead of the previous version.
Phi-4-mini-flash-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 64K token context length.
We introduce Kimi-Dev-72B, our new open-source coding LLM for software engineering tasks. Kimi-Dev-72B achieves a new state-of-the-art on SWE-bench Verified among open-source models.
Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 70B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.
Note for most users: This is an intermediate checkpoint from our post-training pipeline. Most users should use Poro 2 8B Instruct instead, which includes an additional round of Direct Preference Optimization (DPO) for improved response quality and alignment. This SFT-only model is primarily intended for researchers interested in studying the effects of different post-training techniques.
gemma 3n E2B it is een multimodaal taalmodel van Google met 5.4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
⚠️ PREVIEW RELEASE: This is a preview version of EuroMoE-2.6B-A0.6B-Instruct-Preview. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroLLM-2.6B-A0.6-2512, the pre-trained model for EuroLLM-2.6B-A0.6-2512-Instruct.
This is the model card of a 🤗 transformers model that has been pushed on the Hub. This model card has been automatically generated.
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-1.7B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
⚠️ PREVIEW RELEASE: This is a preview version of EuroVLM-9B. The model is still under development and may have limitations in performance and stability. Use with caution in production environments.
This is the model card for EuroLLM-9B-2512, an improved version of utter-project/EuroLLM-9B. In comparison with the previous version, this version includes the long-context extension phase from utter-project/EuroLLM-22B.
Model Summary: Granite Guardian 3.3 8b is a specialized Granite 3.3 8B model designed to judge if the input prompts and the output responses of an LLM based system meet specified criteria. The model comes pre-baked with certain criteria including but not limited to: jailbreak attempts, profanity, and hallucinations related to tool calls and retrieval augmented generation in agent-based systems. Additionally, the model also allows users to bring their own criteria and tailor the judging behavior to specific use-cases.
gemma 3n E4B it is een multimodaal taalmodel van Google met 7.8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
This model was introduced in the paper GUI-Actor: Coordinate-Free Visual Grounding for GUI Agents. It is developed based on UI-TARS-2B-SFT and is designed to predict the correctness of an action position given a language instruction. This model is well-suited for GUI-Actor, as its attention map effectively provides diverse candidates for verification with only a single inference.
The MediPhi Model Collection comprises 7 small language models of 3.8B parameters from the base model Phi-3.5-mini-instruct specialized in the medical and clinical domains. The collection is designed in a modular fashion. Five MediPhi experts are fine-tuned on various medical corpora (i.e. PubMed commercial, Medical Wikipedia, Medical Guidelines, Medical Coding, and open-source clinical documents) and merged back with the SLERP method in their base model to conserve general abilities. One model combined all five experts into one general expert with the multi-model merging method BreadCrumbs. F
The DeepSeek R1 model has undergone a minor version upgrade, with the current version being DeepSeek-R1-0528. In the latest update, DeepSeek R1 has significantly improved its depth of reasoning and inference capabilities by leveraging increased computational resources and introducing algorithmic optimization mechanisms during post-training. The model has demonstrated outstanding performance across various benchmark evaluations, including mathematics, programming, and general logic. Its overall performance is now approaching that of leading models, such as O3 and Gemini 2.5 Pro.
Poro 2 70B Base is a 70B parameter decoder-only transformer created through continued pretraining of Llama 3.1 70B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.
Poro 2 8B Base is an 8B parameter decoder-only transformer created through continued pretraining of Llama 3.1 8B to add Finnish language capabilities. It was trained on 165B tokens using a carefully balanced mix of Finnish, English, code, and math data. Poro 2 is a fully open source model and is made available under the Llama 3.1 Community License.
Poro 2 70B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 70B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.
Poro 2 8B Instruct is an instruction-following chatbot model created through supervised fine-tuning (SFT) and Direct Preference Optimization (DPO) of the Poro 2 8B Base model. This model is designed for conversational AI applications and instruction following in both Finnish and English. It was trained on a carefully curated mix of English and Finnish instruction data, followed by preference tuning to improve response quality.
[!WARNING] WARNING: This model has been deprecated and is no longer recommended. For the latest model, please visit: https://huggingface.co/BSC-LT/Salamandra-VL-7B-2512
medgemma 27b text it is een open-source taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
medgemma 4b it is een multimodaal taalmodel van Google met 4.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Granite Docling 258M builds upon the Idefics3 architecture, but introduces two key modifications: it replaces the vision encoder with siglip2-base-patch16-512 and substitutes the language model with a Granite 165M LLM. Try out our Granite-Docling-258 demo today.
NextCoder: Robust Adaptation of Code LMs to Diverse Code Edits (ICML'2025)
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Model Summary: Granite-4-Tiny-Preview is a 7B parameter fine-grained hybrid mixture-of-experts (MoE) instruct model fine-tuned from Granite-4.0-Tiny-Base-Preview using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised fine-tuning, and model alignment using reinforcement learning.
Phi-4-mini-reasoning is a lightweight open model built upon synthetic data with a focus on high-quality, reasoning dense data further finetuned for more advanced math reasoning capabilities. The model belongs to the Phi-4 model family and supports 128K token context length.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Building upon extensive advancements in training data, model architecture, and optimization techniques, Qwen3 delivers the following key improvements over the previously released Qwen2.5:
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 32B is een open-source taalmodel van Qwen met 32B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 14B is een open-source taalmodel van Qwen met 14B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Qwen3 1.7B is een open-source taalmodel van Qwen met 1.7B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 is the latest generation of large language models in Qwen series, offering a comprehensive suite of dense and mixture-of-experts (MoE) models. Built upon extensive training, Qwen3 delivers groundbreaking advancements in reasoning, instruction-following, agent capabilities, and multilingual support, with the following key features:
Llama Guard 4 12B is een multimodaal taalmodel van Meta met 12B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
NVIDIA Cosmos Reason – an open, customizable, 7B-parameter reasoning vision language model (VLM) for physical AI and robotics - enables robots and vision AI agents to reason like humans, using prior knowledge, physics understanding and common sense to understand and act in the real world. This model understands space, time, and fundamental physics, and can serve as a planning model to reason what steps an embodied agent might take next.
[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).
Eagle2.5 8B is een multimodaal taalmodel van NVIDIA met 8.1B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
[!IMPORTANT] To fully take advantage of the model's capabilities, inference must use temperature=0.8, topk=50, topp=0.95, and dosample=True. For more complex queries, set maxnewtokens=32768 to allow for longer chain-of-thought (CoT).
Model Summary: Granite-3.3-8B-Instruct is a 8-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-8B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on
Model Summary: Granite-3.3-2B-Instruct is a 2-billion parameter 128K context length language model fine-tuned for improved reasoning and instruction-following capabilities. Built on top of Granite-3.3-2B-Base, the model delivers significant gains on benchmarks for measuring generic performance including AlpacaEval-2.0 and Arena-Hard, and improvements in mathematics, coding, and instruction following. It supports structured reasoning through \<think\\<\/think\ and \<response\\<\/response\ tags, providing clear separation between internal thoughts and final outputs. The model has been trained on
Granite-3.3-2B-Base is a decoder-only language model with a 128K token context window. It improves upon Granite-3.1-2B-Base by adding support for Fill-in-the-Middle (FIM) using specialized tokens, enabling the model to generate content conditioned on both prefix and suffix. This makes it well-suited for code completion tasks.
[!Warning] This model has a new version: Kimi-VL-A3B-Thinking-2506. Please consider using the new 2506 version for better abilties on general visual understanding, reasoning, video and agent scenarios.
We present Kimi-VL, an efficient open-source Mixture-of-Experts (MoE) vision-language model (VLM) that offers advanced multimodal reasoning, long-context understanding, and strong agent capabilities—all while activating only 2.8B parameters in its language decoder (Kimi-VL-A3B).
gemma 3 12b it qat q4 0 unquantized is een multimodaal taalmodel van Google met 12B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i
The GLM family welcomes a new generation of open-source models, the GLM-4-32B-0414 series, featuring 32 billion parameters. Its performance is comparable to OpenAI's GPT series and DeepSeek's V3/R1 series, and it supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including a large amount of reasoning-type synthetic data, laying the foundation for subsequent reinforcement learning extensions. In the post-training stage, in addition to human preference alignment for dialogue scenarios, we also enhanced the model's performance i
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
The GLM family welcomes new members, the GLM-4-32B-0414 series models, featuring 32 billion parameters. Its performance is comparable to OpenAI’s GPT series and DeepSeek’s V3/R1 series. It also supports very user-friendly local deployment features. GLM-4-32B-Base-0414 was pre-trained on 15T of high-quality data, including substantial reasoning-type synthetic data. This lays the foundation for subsequent reinforcement learning extensions. In the post-training stage, we employed human preference alignment for dialogue scenarios. Additionally, using techniques like rejection sampling and reinforc
Llama 4 Maverick 17B 128E is een open-source taalmodel van Meta met 17B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Scout 17B 16E is een multimodaal taalmodel van Meta met 109B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Scout 17B 16E Instruct is een multimodaal taalmodel van Meta met 109B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Maverick 17B 128E Instruct is een open-source taalmodel van Meta met 17B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Maverick 17B 128E Instruct FP8 is een open-source taalmodel van Meta met 17B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
DeepSeek-V3-0324 demonstrates notable improvements over its predecessor, DeepSeek-V3, in several key aspects.
txgemma 2b predict is een open-source taalmodel van Google met 2.6B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
salamandra 7b instruct tools is een open-source taalmodel van BSC-LT met 7.8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama-3.3-Nemotron-Super-49B-v1 is a large language model (LLM) which is a derivative of Meta Llama-3.3-70B-Instruct (AKA the reference model). It is a reasoning model that is post trained for reasoning, human chat preferences, and tasks, such as RAG and tool calling. The model supports a context length of 128K tokens.
gemma 3 1b it is een open-source taalmodel van Google met 1B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 12b it is een multimodaal taalmodel van Google met 12B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 12b pt is een multimodaal taalmodel van Google met 12B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 27b pt is een multimodaal taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates
- Weight Decay: Critical for scaling to larger models - Consistent RMS Updates: Enforcing a consistent root mean square on model updates
gemma 3 1b pt is een open-source taalmodel van Google met 1B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 4b it is een multimodaal taalmodel van Google met 4.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 4b pt is een multimodaal taalmodel van Google met 4.3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
🎉Phi-4: [mini-reasoning | reasoning] | [multimodal-instruct | onnx]; [mini-instruct | onnx]
Model Summary: Granite-3.2-2B-Instruct is an 2-billion-parameter, long-context AI model fine-tuned for thinking capabilities. Built on top of Granite-3.1-2B-Instruct, it has been trained using a mix of permissively licensed open-source datasets and internally generated synthetic data designed for reasoning tasks. The model allows controllability of its thinking capability, ensuring it is applied only when required.
Model Summary: granite-vision-3.2-2b is a compact and efficient vision-language model, specifically designed for visual document understanding, enabling automated content extraction from tables, charts, infographics, plots, diagrams, and more. The model was trained on a meticulously curated instruction-following dataset, comprising diverse public datasets and synthetic datasets tailored to support a wide range of document understanding and general image tasks. It was trained by fine-tuning a Granite large language model with both image and text modalities.
In the past five months since Qwen2-VL’s release, numerous developers have built new models on the Qwen2-VL vision-language models, providing us with valuable feedback. During this period, we focused on building more useful vision-language models. Today, we are excited to introduce the latest addition to the Qwen family: Qwen2.5-VL.
--- license: other licensename: qwen licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-72B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
--- license: apache-2.0 language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
--- licensename: qwen-research licenselink: https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct/blob/main/LICENSE language: - en pipelinetag: image-text-to-text tags: - multimodal libraryname: transformers ---
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
DeepSeek R1 Distill Llama 8B is een open-source taalmodel van DeepSeek met 8B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
We introduce our first-generation reasoning models, DeepSeek-R1-Zero and DeepSeek-R1. DeepSeek-R1-Zero, a model trained via large-scale reinforcement learning (RL) without supervised fine-tuning (SFT) as a preliminary step, demonstrated remarkable performance on reasoning. With RL, DeepSeek-R1-Zero naturally emerged with numerous powerful and interesting reasoning behaviors. However, DeepSeek-R1-Zero encounters challenges such as endless repetition, poor readability, and language mixing. To address these issues and further enhance reasoning performance, we introduce DeepSeek-R1, which incorpor
If you are using the weights from this repository, please update to
We are thrilled to release our latest Eagle2 series Vision-Language Model. Open-source Vision-Language Models (VLMs) have made significant strides in narrowing the gap with proprietary models. However, critical details about data strategies and implementation are often missing, limiting reproducibility and innovation. In this project, we focus on VLM post-training from a data-centric perspective, sharing insights into building effective data strategies from scratch. By combining these strategies with robust training recipes and model design, we introduce Eagle2, a family of performant VLMs. Ou
We are thrilled to release our latest Eagle2 series Vision-Language Model. Open-source Vision-Language Models (VLMs) have made significant strides in narrowing the gap with proprietary models. However, critical details about data strategies and implementation are often missing, limiting reproducibility and innovation. In this project, we focus on VLM post-training from a data-centric perspective, sharing insights into building effective data strategies from scratch. By combining these strategies with robust training recipes and model design, we introduce Eagle2, a family of performant VLMs. Ou
We present DeepSeek-V3, a strong Mixture-of-Experts (MoE) language model with 671B total parameters with 37B activated for each token. To achieve efficient inference and cost-effective training, DeepSeek-V3 adopts Multi-head Latent Attention (MLA) and DeepSeekMoE architectures, which were thoroughly validated in DeepSeek-V2. Furthermore, DeepSeek-V3 pioneers an auxiliary-loss-free strategy for load balancing and sets a multi-token prediction training objective for stronger performance. We pre-train DeepSeek-V3 on 14.8 trillion diverse and high-quality tokens, followed by Supervised Fine-Tuning
The CogAgent-9B-20241220 model is based on GLM-4V-9B, a bilingual open-source VLM base model. Through data collection and optimization, multi-stage training, and strategy improvements, CogAgent-9B-20241220 achieves significant advancements in GUI perception, inference prediction accuracy, action space completeness, and task generalizability. The model supports bilingual (Chinese and English) interaction with both screenshots and language input.
Our training data is an extension of the data used for Phi-3 and includes a wide variety of sources from:
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
[!WARNING] WARNING: This is a base language model that has not undergone instruction tuning or alignment with human preferences. As a result, it may generate outputs that are inappropriate, misleading, biased, or unsafe. These risks can be mitigated through additional post-training stages, which is strongly recommended before deployment in any production system, especially for high-stakes applications.
Model Summary: Granite-3.1-1B-A400M-Instruct is a 1B parameter long-context instruct model finetuned from Granite-3.1-1B-A400M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-3B-A800M-Instruct is a 3B parameter long-context instruct model finetuned from Granite-3.1-3B-A800M-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-2B-Instruct is a 2B parameter long-context instruct model finetuned from Granite-3.1-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.1-8B-Instruct is a 8B parameter long-context instruct model finetuned from Granite-3.1-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets tailored for solving long context problems. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Install the transformers library from the source code:
Install the transformers library from the source code:
EuroLLM 9B Instruct is een open-source taalmodel van Utter-project met 9.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
EuroLLM 9B is een open-source taalmodel van Utter-project met 9.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma2 3b pt 448 is een multimodaal taalmodel van Google met 3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma2 3b pt 224 is een multimodaal taalmodel van Google met 3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma2 10b mix 448 is een multimodaal taalmodel van Google met 9.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma2 3b ft docci 448 is een multimodaal taalmodel van Google met 3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma2 3b mix 224 is een multimodaal taalmodel van Google met 3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Install the transformers library from the source code:
Install the transformers library from the source code:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 7B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.
Salamandra is a highly multilingual model pre-trained from scratch that comes in three different sizes — 2B, 7B and 40B parameters — with their respective base and instruction-tuned variants. This model card corresponds to the 2B instructed version specific for AinaHack, an event launched by Generalitat de Catalunya to create AI tools for the Catalan administration.
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
This model is the fp8-quantized version of Salamandra-2b-instruct.
This model is the gptq-quantized version of Salamandra-2b for speculative decoding.
This model is the gptq-quantized version of Salamandra-2b-instruct for speculative decoding.
This model is the gptq-quantized version of Salamandra-7b-instruct for speculative decoding.
This model is the gptq-quantized version of Salamandra-7b for speculative decoding.
This model is the fp8-quantized version of Salamandra-7b.
This model is the fp8-quantized version of Salamandra-2b.
This model is the fp8-quantized version of Salamandra-7b-instruct.
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
If you are using the weights from this repository, please update to
If you are using the weights from this repository, please update to
Granite Guardian 3.0 2B is a fine-tuned Granite 3.0 2B Instruct model designed to detect risks in prompts and responses. It can help with risk detection along many key dimensions catalogued in the IBM AI Risk Atlas. It is trained on unique data comprising human annotations and synthetic data informed by internal red-teaming. It outperforms other open-source models in the same space on standard benchmarks.
This model hub includes a finetuned version of YOLOv8 and a finetuned BLIP-2 model on the above dataset respectively. For more details of the models used and finetuning, please refer to the paper.
Model Summary: Granite-3.0-1B-A400M-Base is a decoder-only language model to support a variety of text-to-text generation tasks. It is trained from scratch following a two-stage training strategy. In the first stage, it is trained on 8 trillion tokens sourced from diverse domains. During the second stage, it is further trained on 2 trillion tokens using a carefully curated mix of high-quality data, aiming to enhance its performance on specific tasks.
Model Summary: Granite-3.0-3B-A800M-Instruct is a 3B parameter model finetuned from Granite-3.0-3B-A800M-Base-4K using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.0-8B-Instruct is a 8B parameter model finetuned from Granite-3.0-8B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
Model Summary: Granite-3.0-2B-Instruct is a 2B parameter model finetuned from Granite-3.0-2B-Base using a combination of open source instruction datasets with permissive license and internally collected synthetic datasets. This model is developed using a diverse set of techniques with a structured chat format, including supervised finetuning, model alignment using reinforcement learning, and model merging.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
This repository contains the model described in Salamandra Technical Report.
gemma 2 2b jpn it is een open-source taalmodel van Google met 2.6B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
--- license: apache-2.0 datasets: - projecte-aina/RAGMultilingual language: - es - en - ca libraryname: transformers ---
- Developed by: Fraunhofer, Forschungszentrum Jülich, TU Dresden, DFKI - Funded by: German Federal Ministry of Economics and Climate Protection (BMWK) in the context of the OpenGPT-X project - Model type: Transformer based decoder-only model - Language(s) (NLP): bg, cs, da, de, el, en, es, et, fi, fr, ga, hr, hu, it, lt, lv, mt, nl, pl, pt, ro, sk, sl, sv - Shared by: OpenGPT-X
--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---
Llama Guard 3 1B is een open-source taalmodel van Meta met 1.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama Guard 3 11B Vision is een multimodaal taalmodel van Meta met 11B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 90B Vision Instruct is een multimodaal taalmodel van Meta met 89B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 90B Vision is een multimodaal taalmodel van Meta met 89B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 11B Vision Instruct is een multimodaal taalmodel van Meta met 11B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 11B Vision is een open-source taalmodel van Meta met 11B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 3B is een open-source taalmodel van Meta met 3B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 3B Instruct is een open-source taalmodel van Meta met 3.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 1B Instruct is een open-source taalmodel van Meta met 1.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.2 1B is een open-source taalmodel van Meta met 1B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
--- license: apache-2.0 language: - en - ca - es basemodel: - projecte-aina/FLOR-6.3B pipelinetag: text-generation libraryname: transformers ---
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5-Coder is the latest series of Code-Specific Qwen large language models (formerly known as CodeQwen). As of now, Qwen2.5-Coder has covered six mainstream model sizes, 0.5, 1.5, 3, 7, 14, 32 billion parameters, to meet the needs of different developers. Qwen2.5-Coder brings the following improvements upon CodeQwen1.5:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Qwen2.5 is the latest series of Qwen large language models. For Qwen2.5, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters. Qwen2.5 brings the following improvements upon Qwen2:
Nemotron-Mini-4B-Instruct is a model for generating responses for roleplaying, retrieval augmented generation, and function calling. It is a small language model (SLM) optimized through distillation, pruning and quantization for speed and on-device deployment. It is a fine-tuned version of nvidia/Minitron-4B-Base, which was pruned and distilled from Nemotron-4 15B using our LLM compression technique. This instruct model is optimized for roleplay, RAG QA, and function calling in English. It supports a context length of 4,096 tokens. This model is ready for commercial use.
gemma 7b aps it is een open-source taalmodel van Google met 8.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
2b Version of model Sxxxxxx, without last epoch, instructed with baseline dataset including RAGMultilingual
This model is ready for commercial and non-commercial use. <br
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.
We're excited to unveil Qwen2-VL, the latest iteration of our Qwen-VL model, representing nearly a year of innovation.
Phi-3.5-MoE is a lightweight, state-of-the-art open model built upon datasets used for Phi-3 - synthetic data and filtered publicly available documents - with a focus on very high-quality, reasoning dense data. The model supports multilingual and comes with 128K context length (in tokens). The model underwent a rigorous enhancement process, incorporating supervised fine-tuning, proximal policy optimization, and direct preference optimization to ensure precise instruction adherence and robust safety measures.
🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]
LongWriter-glm4-9b is trained based on glm-4-9b, and is capable of generating 10,000+ words at once.
LongWriter-llama3.1-8b is trained based on Meta-Llama-3.1-8B, and is capable of generating 10,000+ words at once.
This instructed model uses a chat template that must be adhered to the input for conversational use. The easiest way to apply it is using the tokenizer's built-in chat template, as shown in the following snippet.
This is the model card for the first instruction tuned model of the EuroLLM series: EuroLLM-1.7B-Instruct. You can also check the pre-trained version: EuroLLM-1.7B.
This is the model card for the first pre-trained model of the EuroLLM series: EuroLLM-1.7B. You can also check the instruction tuned version: EuroLLM-1.7B-Instruct.
experimental7b rag instruct is een open-source taalmodel van BSC-LT met 7.8B parameters en een contextvenster van 8K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
AWQ quantized version of google/gemma-2b.
Llama Guard 3 8B is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama Guard 3 8B INT8 is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 405B FP8 is een open-source taalmodel van Meta met 405B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 405B Instruct FP8 is een open-source taalmodel van Meta met 405B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
experimental7b rag is een open-source taalmodel van BSC-LT met 7.8B parameters en een contextvenster van 8K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
shieldgemma 9b is een open-source taalmodel van Google met 9.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
shieldgemma 2b is een open-source taalmodel van Google met 2.6B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 405B Instruct is een open-source taalmodel van Meta met 405B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 70B Instruct is een open-source taalmodel van Meta met 71B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2 2b it is een open-source taalmodel van Google met 2.6B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2 2b is een open-source taalmodel van Google met 2.6B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 405B is een open-source taalmodel van Meta met 405B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 70B is een open-source taalmodel van Meta met 70B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.1 8B is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
We introduce CodeGeeX4-ALL-9B, the open-source version of the latest CodeGeeX4 model series. It is a multilingual code generation model continually trained on the GLM-4-9B, significantly enhancing its code generation capabilities. Using a single CodeGeeX4-ALL-9B model, it can support comprehensive functions such as code completion and generation, code interpreter, web search, function call, repository-level code Q&A, covering various scenarios of software development. CodeGeeX4-ALL-9B has achieved highly competitive performance on public benchmarks, such as BigCodeBench and NaturalCodeBench. I
gemma 2 9b is een open-source taalmodel van Google met 9.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2 9b it is een open-source taalmodel van Google met 9.2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2 27b is een open-source taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2 27b it is een open-source taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.
In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.
In standard benchmark evaluations, DeepSeek-Coder-V2 achieves superior performance compared to closed-source models such as GPT4-Turbo, Claude 3 Opus, and Gemini 1.5 Pro in coding and math benchmarks. The list of supported programming languages can be found here.
2024/08/12, 本仓库代码已更新并使用 transformers=4.44.0, 请及时更新依赖。
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the instruction-tuned 1.5B Qwen2 model.
Qwen2 is the new series of Qwen large language models. For Qwen2, we release a number of base language models and instruction-tuned language models ranging from 0.5 to 72 billion parameters, including a Mixture-of-Experts model. This repo contains the 0.5B Qwen2 base language model.
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:
Last week, the release and buzz around DeepSeek-V2 have ignited widespread interest in MLA (Multi-head Latent Attention)! Many in the community suggested open-sourcing a smaller MoE model for in-depth research. And now DeepSeek-V2-Lite comes out:
paligemma 3b ft cococap 448 is een multimodaal taalmodel van Google met 2.9B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma 3b pt 448 is een multimodaal taalmodel van Google met 2.9B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma 3b mix 224 is een multimodaal taalmodel van Google met 2.9B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
paligemma 3b pt 224 is een multimodaal taalmodel van Google met 2.9B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
New applications/projects should use the latest mainline Granite language model family, whose code capabilities supercede this model. This model is being made available strictly for historical/scientific purposes. Please see our Granite Collections for the latest Granite releases.
🎉Phi-4: [multimodal-instruct | onnx]; [mini-instruct | onnx]
🎉 Phi-3.5: [[mini-instruct]](https://huggingface.co/microsoft/Phi-3.5-mini-instruct); [[MoE-instruct]](https://huggingface.co/microsoft/Phi-3.5-MoE-instruct) ; [[vision-instruct]](https://huggingface.co/microsoft/Phi-3.5-vision-instruct)
Due to the constraints of HuggingFace, the open-source code currently experiences slower performance than our internal codebase when running on GPUs with Huggingface. To facilitate the efficient execution of our model, we offer a dedicated vllm solution that optimizes performance for running our model effectively.
Meta Llama Guard 2 8B is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Meta Llama 3 8B is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Meta Llama 3 8B Instruct is een open-source taalmodel van Meta met 8B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Meta Llama 3 70B Instruct is een open-source taalmodel van Meta met 71B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Meta Llama 3 70B is een open-source taalmodel van Meta met 71B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Poro 34b chat is a chat-tuned version of Poro 34B trained to follow instructions in both Finnish and English. Quantized versions are available on Poro 34B-chat-GGUF.
gemma 1.1 2b it is een open-source taalmodel van Google met 2.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 1.1 7b it is een open-source taalmodel van Google met 8.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
codegemma 2b is een open-source taalmodel van Google met 2.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 34b Instruct hf is een open-source taalmodel van Meta met 34B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 70b Instruct hf is een open-source taalmodel van Meta met 69B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 13b Instruct hf is een open-source taalmodel van Meta met 13B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 13b hf is een open-source taalmodel van Meta met 13B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 7b Instruct hf is een open-source taalmodel van Meta met 6.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeLlama 7b hf is een open-source taalmodel van Meta met 6.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Viking 33B is a 33B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1300B billion as of this release). Viking 33B is a fully open source model and is made available under the Apache 2.0 License.
Viking 13B is a 13B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It is being trained on 2 trillion tokens (1.3 trillion as of this release). Viking 13B is a fully open source model and is made available under the Apache 2.0 License.
Viking 7B is a 7B parameter decoder-only transformer pretrained on Finnish, English, Swedish, Danish, Norwegian, Icelandic and code. It has been trained on 2 trillion tokens. Viking 7B is a fully open source model and is made available under the Apache 2.0 License.
gemma 7b it is een open-source taalmodel van Google met 8.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 7b is een open-source taalmodel van Google met 8.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2b it is een open-source taalmodel van Google met 2.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 2b is een open-source taalmodel van Google met 2.5B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
❗❗❗ Please use chain-of-thought prompt to test DeepSeekMath-Instruct and DeepSeekMath-RL:
deepseek coder 7b instruct v1.5 is een open-source taalmodel van DeepSeek met 7B parameters en een contextvenster van 4K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
python import torch from transformers import AutoTokenizer, AutoModelForCausalLM, GenerationConfig
modelname = "deepseek-ai/deepseek-moe-16b-base" tokenizer = AutoTokenizer.frompretrained(modelname) model = AutoModelForCausalLM.frompretrained(modelname, torchdtype=torch.bfloat16, devicemap="auto") model.generationconfig = GenerationConfig.frompretrained(modelname) model.generationconfig.padtokenid = model.generationconfig.eostokenid
Phi-2 is a Transformer with 2.7 billion parameters. It was trained using the same data sources as Phi-1.5, augmented with a new data source that consists of various NLP synthetic texts and filtered websites (for safety and educational value). When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-2 showcased a nearly state-of-the-art performance among models with less than 13 billion parameters.
LlamaGuard 7b is een open-source taalmodel van Meta met 6.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Introducing DeepSeek LLM, an advanced language model comprising 67 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
Introducing DeepSeek LLM, an advanced language model comprising 7 billion parameters. It has been trained from scratch on a vast dataset of 2 trillion tokens in both English and Chinese. In order to foster research, we have made DeepSeek LLM 7B/67B Base and DeepSeek LLM 7B/67B Chat open source for the research community.
- Repository: https://github.com/thu-coai/BPO - Paper: https://arxiv.org/abs/2311.04155 - Data: https://huggingface.co/datasets/THUDM/BPO
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Deepseek Coder is composed of a series of code language models, each trained from scratch on 2T tokens, with a composition of 87% code and 13% natural language in both English and Chinese. We provide various sizes of the code model, ranging from 1B to 33B versions. Each model is pre-trained on project-level code corpus by employing a window size of 16K and a extra fill-in-the-blank task, to support project-level code completion and infilling. For coding capabilities, Deepseek Coder achieves state-of-the-art performance among open-source code models on multiple programming languages and various
Poro is a 34B parameter decoder-only transformer pretrained on Finnish, English and code. It was trained on 1 trillion tokens. Poro is a fully open source model and is made available under the Apache 2.0 License.
py from mistralcommon.tokens.tokenizers.mistral import MistralTokenizer from mistralcommon.protocol.instruct.messages import UserMessage from mistralcommon.protocol.instruct.request import ChatCompletionRequest
The Mistral-7B-v0.1 Large Language Model (LLM) is a pretrained generative text model with 7 billion parameters. Mistral-7B-v0.1 outperforms Llama 2 13B on all benchmarks we tested.
The language model Phi-1 is a Transformer with 1.3 billion parameters, specialized for basic Python coding. Its training involved a variety of data sources, including subsets of Python codes from The Stack v1.2, Q&A content from StackOverflow, competition code from codecontests, and synthetic Python textbooks and exercises generated by gpt-3.5-turbo-0301. Even though the model and the datasets are relatively small compared to contemporary Large Language Models (LLMs), Phi-1 has demonstrated an impressive accuracy rate exceeding 50% on the simple Python coding benchmark, HumanEval.
The language model Phi-1.5 is a Transformer with 1.3 billion parameters. It was trained using the same data sources as phi-1, augmented with a new data source that consists of various NLP synthetic texts. When assessed against benchmarks testing common sense, language understanding, and logical reasoning, Phi-1.5 demonstrates a nearly state-of-the-art performance among models with less than 10 billion parameters.
Llama 2 70b chat hf is een open-source taalmodel van Meta met 69B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 2 7b chat hf is een open-source taalmodel van Meta met 6.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 2 7b hf is een open-source taalmodel van Meta met 6.7B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 2 13b hf is een open-source taalmodel van Meta met 13B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 2 13b chat hf is een open-source taalmodel van Meta met 13B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 2 70b hf is een open-source taalmodel van Meta met 69B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
CodeGPT small py is een open-source taalmodel van Microsoft, gehost op Europese GPU's via een OpenAI-compatibele API.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
DialoGPT is a SOTA large-scale pretrained dialogue response generation model for multiturn conversations. The human evaluation results indicate that the response generated from DialoGPT is comparable to human response quality under a single-turn conversation Turing test. The model is trained on 147M multi-turn dialogue from Reddit discussion thread.
Gemma 3 1B is een open-source taalmodel van Google met 1B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 30B A3B Instruct 2507 is een open-source taalmodel van Qwen met 30B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 4B Thinking 2507 is een open-source taalmodel van Qwen met 4B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3 12b is een open-source taalmodel van Google met 12B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
functiongemma 270m is een open-source taalmodel van Google, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek V3 (685B MoE) is een open-source taalmodel van DeepSeek met 685B parameters en een contextvenster van 164K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 VL 30B A3B is een open-source taalmodel van Qwen met 30B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek Coder V2 Lite (16B) is een open-source taalmodel van DeepSeek met 16B parameters en een contextvenster van 164K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
DeepSeek R1 Distill 1.5B is een open-source taalmodel van DeepSeek met 1.5B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 VL 7B is een open-source taalmodel van Qwen met 7B parameters en een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3 VL 8B Thinking is een open-source taalmodel van Qwen met 8B parameters en een contextvenster van 262K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
deepseek coder 6.7b is een open-source taalmodel van DeepSeek met 6.7B parameters en een contextvenster van 16K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen3Guard Gen 0.6B is een open-source taalmodel van Qwen met 0.6B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
translategemma 4b is een open-source taalmodel van Google met 4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 VL 72B is een open-source taalmodel van Qwen met 72B parameters en een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Mistral Nemo 12B is een open-source taalmodel van Mistral met 12B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 3.3 70B is een open-source taalmodel van Meta met 70B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Maverick (17Bx128E) is een open-source taalmodel van Meta met 17B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Llama 4 Scout (17Bx16E) is een open-source taalmodel van Meta met 17B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Sovereign EU model fine-tuned by HostYourAI on dutch-clean.
Sovereign EU model fine-tuned by HostYourAI on loes-xl-pre.
Sovereign EU model fine-tuned by HostYourAI on loes-large-v1.
medgemma 27b is een open-source taalmodel van Google met 27B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
medgemma 4b is een open-source taalmodel van Google met 4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Mistral Medium 3.5 is een open-source taalmodel van Mistral-medium-3.5-128b met een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 VL 3B is een open-source taalmodel van Qwen met 3B parameters en een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Phi 4 Mini (3.8B) is een open-source taalmodel van Microsoft met 3.8B parameters en een contextvenster van 131K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
gemma 3n E2B is een open-source taalmodel van Google met 2B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 2.5 14B is een open-source taalmodel van Qwen met 14B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 2.5 72B is een open-source taalmodel van Qwen met 72B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 2.5 Coder 1.5B is een open-source taalmodel van Qwen met 1.5B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 2.5 Coder 32B is een open-source taalmodel van Qwen met 32B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 2.5 Coder 7B is een open-source taalmodel van Qwen met 7B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 3 0.6B is een open-source taalmodel van Qwen met 0.6B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen 3 4B is een open-source taalmodel van Qwen met 4B parameters en een contextvenster van 41K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Gemma 3 4B is een open-source taalmodel van Google met 4B parameters, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 32B is een open-source taalmodel van Qwen met 32B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 7B is een open-source taalmodel van Qwen met 7B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 Coder 14B is een open-source taalmodel van Qwen met 14B parameters en een contextvenster van 33K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Qwen2.5 VL 32B is een open-source taalmodel van Qwen met 32B parameters en een contextvenster van 128K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Phi 4 (14B) is een open-source taalmodel van Microsoft met 14B parameters en een contextvenster van 16K tokens, gehost op Europese GPU's via een OpenAI-compatibele API.
Geen creditcard nodig. Betaal naar gebruik, stop wanneer je wilt.
Begin vandaag gratis met hosten