Last updated: 10/6/2026Last verified: 2026-10-03

Kolibri

Kolibri Review

2
0

Kolibri is Aleph Alpha's Apache-2.0 open-weight German-English mixture-of-experts reasoning model with 78B total and 3.46B active parameters, up to 1M-token context, tool calling, and self-hosted vLLM deployment.

Free

What is Kolibri?

Kolibri is Aleph Alpha’s Apache-2.0 open-weight German-English mixture-of-experts reasoning model. It has 78 billion total parameters, with 3.46 billion active parameters, and supports reasoning mode, tool calling, and context lengths validated up to 1,048,576 tokens. Developers can self-host it through an official vLLM plugin with OpenAI-compatible serving.

Key Features

  • Mixture-of-experts architecture with 78B total parameters and 3.46B active parameters
  • Native support for German and English prompts and responses
  • Reasoning mode for tasks that benefit from structured problem-solving
  • Tool calling support for connecting model outputs with external functions or services
  • Context length validated up to 1,048,576 tokens
  • Apache-2.0 open-weight licensing
  • Self-hosted deployment through the official vLLM plugin
  • OpenAI-compatible serving for applications built around compatible APIs

Best For

German-English chat and assistant applicationsLong-context analysis of large documents or codebasesDevelopers building self-hosted AI applicationsResearch workflows that require an open-weight reasoning modelTool-using agents and API-connected automationCoding applications that can integrate with OpenAI-compatible endpoints

Pricing

Kolibri is listed as free, and its weights use the Apache-2.0 license. Deployment through vLLM is self-hosted, so users should evaluate their own hardware, hosting, and operational requirements separately.

Pros & Cons

Pros

  • Combines German-English language support with reasoning capabilities
  • Uses an open-weight Apache-2.0 model license
  • Offers a validated context length of up to 1,048,576 tokens
  • Supports tool calling for integrations and agent workflows
  • Provides OpenAI-compatible serving through the official vLLM plugin
  • The MoE design uses 3.46B active parameters despite 78B total parameters

Cons

  • Self-hosting through vLLM requires deployment and infrastructure work
  • The 78B total parameter count may make infrastructure planning more involved
  • The model is primarily positioned for German-English use rather than broad multilingual coverage
  • A hosted chat interface or managed API is not specified in the provided information
  • Actual output quality, latency, and hardware requirements should be tested for each workload

Alternatives

DeepSeek-R1

An open-weight reasoning model that can be considered for research, coding, and problem-solving workflows.

Qwen

The Qwen model family offers open-weight language models for chat, coding, research, and self-hosted applications.

Llama

Meta’s Llama model family is a widely used option for developers evaluating open-weight language models and self-hosted deployments.

Mistral

Mistral provides open and downloadable language models that can serve as alternatives for coding, assistant, and research applications.

FAQ

AD

Details

Platform

LinuxAPI

Features

  • 78B-total and 3.46B-active mixture-of-experts architecture
  • Native German and English support with reasoning mode and tool calling
  • Context length validated up to 1,048,576 tokens
  • OpenAI-compatible serving through the official vLLM plugin and Apache-2.0 weights

Languages

ende

Known limitations

  • The full FP8 model has an approximately 78 GB memory footprint and needs multi-GPU or high-memory hardware
  • The official model card recommends contexts of at most 262,144 tokens for serving efficiency and complex tasks
  • Apache-2.0 applies to the published weights and configuration files, not all underlying code or model IP
  • The model is intended for human-reviewed assistance rather than unreviewed autonomous decisions

Rate This Tool

Related Tools