Last updated: 10/6/2026Last verified: 2026-10-01

AA-AgentPerf-Local

AA-AgentPerf-Local Review

6
0

Open-source local AI agent inference benchmark from Artificial Analysis. It replays recorded agent trajectories against OpenAI-compatible model servers to measure throughput and latency on laptops and workstations.

Free

What is AA-AgentPerf-Local?

AA-AgentPerf-Local is an open-source local AI agent inference benchmark from Artificial Analysis. It replays recorded agent trajectories against OpenAI-compatible model servers to measure throughput and latency on laptops and workstations. The benchmark is designed for users who want reproducible performance data for local model serving setups.

Key Features

  • Replays eight recorded agent tasks across 168 model turns with a growing real-world context
  • Measures completion time, prefill speed, decode speed, and time to first token
  • Runs against existing OpenAI-compatible servers
  • Can manage pinned local model recipes for more reproducible testing
  • Publishes reproducible hardware, model, runtime, and serving configurations
  • Focused on laptop and workstation inference performance for agent-style workloads

Best For

Developers benchmarking local LLM serving stacksAI researchers comparing inference performance across hardware and runtimesTeams evaluating workstation or laptop suitability for agentic AI workloadsUsers running OpenAI-compatible local model serversPractitioners who need reproducible latency and throughput measurements

Pricing

AA-AgentPerf-Local is listed as a free, open-source tool. No paid tiers or commercial pricing details were provided in the available tool information.

Pros & Cons

Pros

  • Uses recorded agent trajectories rather than only synthetic single-prompt tests
  • Covers multiple latency and throughput metrics, including time to first token, prefill speed, and decode speed
  • Works with OpenAI-compatible model servers, making it useful across different local serving setups
  • Supports reproducibility by publishing hardware, model, runtime, and serving configurations
  • Targets practical local hardware such as laptops and workstations

Cons

  • Scope is focused on local inference benchmarking rather than general model evaluation
  • Requires access to a compatible local or OpenAI-compatible serving environment
  • Benchmark results may depend heavily on hardware, model choice, runtime, and configuration
  • The provided information does not describe a hosted dashboard, collaboration tools, or enterprise support

Alternatives

MLPerf Inference

A widely used benchmark suite for measuring machine learning inference performance across hardware and software systems.

lm-evaluation-harness

An open-source framework for evaluating language models on a broad set of NLP and reasoning benchmarks.

OpenAI Evals

An evaluation framework for testing model behavior and performance against custom evaluation tasks.

HELM

A language model evaluation benchmark focused on standardized model assessment across multiple scenarios and metrics.

vLLM Benchmarks

Benchmarking utilities associated with vLLM for measuring serving throughput and latency in OpenAI-compatible inference deployments.

FAQ

AD

Details

Platform

macOSWindowsLinux

Features

  • Replays eight recorded agent tasks across 168 model turns with a growing real-world context
  • Measures completion time, prefill speed, decode speed, and time to first token
  • Runs against existing OpenAI-compatible servers or manages pinned local model recipes
  • Publishes reproducible hardware, model, runtime, and serving configurations

Languages

en

Known limitations

  • The default comparable workload requires a 65,536-token context and a local OpenAI-compatible model server
  • Results measure inference speed rather than output quality, and tool execution is skipped by default
  • Managed Windows runs require an NVIDIA GPU; vLLM and SGLang recipes run only on Linux
  • The public hardware and model leaderboard is still limited and is planned to expand over time

Rate This Tool

Related Tools