Last updated: 10/6/2026Last verified: 2026-10-03
What is Kolibri?
Kolibri is Aleph Alpha’s Apache-2.0 open-weight German-English mixture-of-experts reasoning model. It has 78 billion total parameters, with 3.46 billion active parameters, and supports reasoning mode, tool calling, and context lengths validated up to 1,048,576 tokens. Developers can self-host it through an official vLLM plugin with OpenAI-compatible serving.
Key Features
- Mixture-of-experts architecture with 78B total parameters and 3.46B active parameters
- Native support for German and English prompts and responses
- Reasoning mode for tasks that benefit from structured problem-solving
- Tool calling support for connecting model outputs with external functions or services
- Context length validated up to 1,048,576 tokens
- Apache-2.0 open-weight licensing
- Self-hosted deployment through the official vLLM plugin
- OpenAI-compatible serving for applications built around compatible APIs
Best For
Pricing
Kolibri is listed as free, and its weights use the Apache-2.0 license. Deployment through vLLM is self-hosted, so users should evaluate their own hardware, hosting, and operational requirements separately.
Pros & Cons
Pros
- Combines German-English language support with reasoning capabilities
- Uses an open-weight Apache-2.0 model license
- Offers a validated context length of up to 1,048,576 tokens
- Supports tool calling for integrations and agent workflows
- Provides OpenAI-compatible serving through the official vLLM plugin
- The MoE design uses 3.46B active parameters despite 78B total parameters
Cons
- Self-hosting through vLLM requires deployment and infrastructure work
- The 78B total parameter count may make infrastructure planning more involved
- The model is primarily positioned for German-English use rather than broad multilingual coverage
- A hosted chat interface or managed API is not specified in the provided information
- Actual output quality, latency, and hardware requirements should be tested for each workload
Alternatives
An open-weight reasoning model that can be considered for research, coding, and problem-solving workflows.
The Qwen model family offers open-weight language models for chat, coding, research, and self-hosted applications.
Meta’s Llama model family is a widely used option for developers evaluating open-weight language models and self-hosted deployments.
Mistral provides open and downloadable language models that can serve as alternatives for coding, assistant, and research applications.
FAQ
Details
Platform
Features
- 78B-total and 3.46B-active mixture-of-experts architecture
- Native German and English support with reasoning mode and tool calling
- Context length validated up to 1,048,576 tokens
- OpenAI-compatible serving through the official vLLM plugin and Apache-2.0 weights
Languages
Known limitations
- The full FP8 model has an approximately 78 GB memory footprint and needs multi-GPU or high-memory hardware
- The official model card recommends contexts of at most 262,144 tokens for serving efficiency and complex tasks
- Apache-2.0 applies to the published weights and configuration files, not all underlying code or model IP
- The model is intended for human-reviewed assistance rather than unreviewed autonomous decisions






