Last updated: 10/6/2026Last verified: 2026-10-01
What is AA-AgentPerf-Local?
AA-AgentPerf-Local is an open-source local AI agent inference benchmark from Artificial Analysis. It replays recorded agent trajectories against OpenAI-compatible model servers to measure throughput and latency on laptops and workstations. The benchmark is designed for users who want reproducible performance data for local model serving setups.
Key Features
- Replays eight recorded agent tasks across 168 model turns with a growing real-world context
- Measures completion time, prefill speed, decode speed, and time to first token
- Runs against existing OpenAI-compatible servers
- Can manage pinned local model recipes for more reproducible testing
- Publishes reproducible hardware, model, runtime, and serving configurations
- Focused on laptop and workstation inference performance for agent-style workloads
Best For
Pricing
AA-AgentPerf-Local is listed as a free, open-source tool. No paid tiers or commercial pricing details were provided in the available tool information.
Pros & Cons
Pros
- Uses recorded agent trajectories rather than only synthetic single-prompt tests
- Covers multiple latency and throughput metrics, including time to first token, prefill speed, and decode speed
- Works with OpenAI-compatible model servers, making it useful across different local serving setups
- Supports reproducibility by publishing hardware, model, runtime, and serving configurations
- Targets practical local hardware such as laptops and workstations
Cons
- Scope is focused on local inference benchmarking rather than general model evaluation
- Requires access to a compatible local or OpenAI-compatible serving environment
- Benchmark results may depend heavily on hardware, model choice, runtime, and configuration
- The provided information does not describe a hosted dashboard, collaboration tools, or enterprise support
Alternatives
A widely used benchmark suite for measuring machine learning inference performance across hardware and software systems.
An open-source framework for evaluating language models on a broad set of NLP and reasoning benchmarks.
An evaluation framework for testing model behavior and performance against custom evaluation tasks.
A language model evaluation benchmark focused on standardized model assessment across multiple scenarios and metrics.
Benchmarking utilities associated with vLLM for measuring serving throughput and latency in OpenAI-compatible inference deployments.
FAQ
Details
Platform
Features
- Replays eight recorded agent tasks across 168 model turns with a growing real-world context
- Measures completion time, prefill speed, decode speed, and time to first token
- Runs against existing OpenAI-compatible servers or manages pinned local model recipes
- Publishes reproducible hardware, model, runtime, and serving configurations
Languages
Known limitations
- The default comparable workload requires a 65,536-token context and a local OpenAI-compatible model server
- Results measure inference speed rather than output quality, and tool execution is skipped by default
- Managed Windows runs require an NVIDIA GPU; vLLM and SGLang recipes run only on Linux
- The public hardware and model leaderboard is still limited and is planned to expand over time






