Last updated: 10/6/2026Last verified: 2026-08-14
Toast 1 is Mixedbread's agentic search model: it decomposes a query into steps, runs parallel retrievals with "next file search" iteration, inspects sources, and curates evidence before returning. It runs standalone or as a retrieval subagent, delivering frontier search quality at roughly 1/10th the cost and 12x the speed of comparable frontier models.
What is Toast 1?
Toast 1 is Mixedbread's agentic search model designed to improve retrieval quality for AI search, research, and coding workflows. It breaks queries into steps, runs parallel retrieval operations, inspects sources, and curates evidence before returning results. The model can run as a standalone search agent or as a retrieval subagent inside larger frontier-agent workflows.
Key Features
- Decomposes complex queries into smaller retrieval steps before returning results
- Runs up to 8 parallel retrieval calls per round, including semantic search, grep, metadata filtering, and context pruning
- Uses a bounded agentic loop with up to 4 rounds to gather evidence and inspect sources
- Returns a final ranked chunk list designed for retrieval-augmented generation and agent workflows
- Supports standalone operation as a specialized search agent
- Can be used as a retrieval subagent inside tools and coding agents such as Codex-style or OpenCode-style workflows
- Offers a Chat Completions-compatible API with an agentic flag for easier integration into existing stacks
Best For
Pricing
Toast 1 is listed as a paid tool. The provided information does not include specific pricing tiers, usage-based rates, free trial details, or enterprise contract terms, so buyers should check Mixedbread's official website for current pricing.
Pros & Cons
Pros
- Agentic retrieval process can break down complex queries instead of relying on a single search pass
- Parallel retrieval calls may help cover semantic, keyword, metadata, and contextual signals in one workflow
- Bounded loop design gives the retrieval agent a defined maximum number of rounds
- Ranked chunk output is well suited for RAG systems and AI coding assistants
- Chat Completions-compatible API may reduce integration work for teams with existing LLM infrastructure
Cons
- Specific pricing details are not included in the provided information
- Performance claims such as lower cost and faster speed versus comparable frontier models need independent verification
- The review data does not provide details about security, compliance, data retention, or enterprise controls
- Teams may still need to benchmark Toast 1 against their own corpus, retrieval stack, and latency requirements
Alternatives
Perplexity is an AI search and research tool focused on web-based answers with cited sources, making it relevant for users evaluating AI-assisted research workflows.
Exa provides AI-native search APIs for retrieving web content and can be used in research agents, RAG systems, and custom search applications.
Tavily offers search APIs designed for AI agents, including retrieval features aimed at grounding LLM responses with external information.
Sourcegraph Cody is an AI coding assistant that uses code search and repository context, making it relevant for teams focused on code retrieval and developer workflows.
Glean provides enterprise AI search across workplace knowledge sources and is an alternative for organizations prioritizing internal knowledge retrieval.
FAQ
Details
Platform
Features
- Decomposes queries into steps and iterates with up to 8 parallel retrieval calls per round (semantic search, grep, metadata filter, context pruning)
- Bounded agentic loop (4 rounds max) that gathers evidence, checks sources, and returns a final ranked chunk list
- Runs standalone as a specialized search agent or as a retrieval subagent inside frontier agents (e.g. Codex, OpenCode)
- Drop-in for existing stacks: Chat Completions-compatible API with an agentic flag and the same ranked-chunk response shape
Languages
Known limitations
- API-only hosted model — no open weights on HuggingFace and model size is undisclosed as of 2026-08-14
- Launched 2026-08-13 — benchmarks are vendor-reported and independent third-party evaluation is still pending
- Slower than a single search call (extra LLM + retrieval rounds), and the loop is bounded to 4 rounds, limiting very deep retrieval
- Strongest with Mixedbread Search as the backend; other retrieval backends are competitive but not at full strength








