最后更新:10/6/2026最后核验:2026-10-01
AA-AgentPerf-Local 是什么
AA-AgentPerf-Local 是 Artificial Analysis 推出的开源本地 AI Agent 推理性能基准测试工具。它会将已记录的 Agent 运行轨迹回放到兼容 OpenAI API 的模型服务器上,用于衡量笔记本电脑和工作站上的吞吐量与延迟表现。该 Benchmark 面向希望获得可复现本地模型服务性能数据的开发者、研究人员和技术团队。
核心功能
- 回放 8 个已记录的 Agent 任务,覆盖 168 轮模型调用,并逐步引入更接近真实场景的上下文
- 衡量任务完成时间、prefill speed、decode speed 以及 time to first token 等关键指标
- 可直接对接现有 OpenAI-compatible model servers
- 支持管理固定版本的本地模型配置方案,提升测试可复现性
- 公开硬件、模型、runtime 与 serving 配置,便于复现实验结果
- 专注于笔记本和工作站在 Agent 式工作负载下的本地推理性能
最适合
需要评测本地 LLM serving stack 的开发者希望比较不同硬件和 runtime 推理性能的 AI 研究人员评估工作站或笔记本是否适合运行 Agentic AI 工作负载的团队正在使用 OpenAI-compatible 本地模型服务器的用户需要可复现延迟与吞吐量测试结果的 AI 工程实践者
定价
AA-AgentPerf-Local 目前被列为免费、开源工具。根据现有工具信息,未提供付费套餐、商业授权或企业版定价细节。
优缺点
优点
- 使用已记录的 Agent 运行轨迹,而不只是简单的合成单 prompt 测试
- 覆盖多项延迟和吞吐量指标,包括 time to first token、prefill speed 与 decode speed
- 兼容 OpenAI-compatible model servers,适用于多种本地模型 serving 环境
- 通过公开硬件、模型、runtime 和 serving 配置,增强测试结果的可复现性
- 面向笔记本电脑和工作站等实际本地硬件场景,实用性较强
缺点
- 定位主要是本地推理 Benchmark,不适合替代通用模型能力评测
- 需要具备兼容的本地 serving 环境或 OpenAI-compatible 模型服务器
- 测试结果会受到硬件、模型选择、runtime 和配置参数的显著影响
- 现有信息未提及托管 dashboard、团队协作功能或企业级支持
替代工具
MLPerf Inference
广泛使用的机器学习推理性能基准测试套件,可用于衡量不同硬件与软件系统的 inference performance。
lm-evaluation-harness
开源语言模型评测框架,支持在大量 NLP 与推理类 Benchmark 上评估模型能力。
OpenAI Evals
用于针对自定义评测任务测试模型行为与性能的评估框架。
HELM
面向语言模型标准化评估的 Benchmark,覆盖多种应用场景和评测指标。
vLLM Benchmarks
与 vLLM 相关的 Benchmark 工具,可用于测量 OpenAI-compatible 推理部署中的 serving 吞吐量和延迟。
常见问题
广告
详情
平台
macOSWindowsLinux
功能特性
- Replays eight recorded agent tasks across 168 model turns with a growing real-world context
- Measures completion time, prefill speed, decode speed, and time to first token
- Runs against existing OpenAI-compatible servers or manages pinned local model recipes
- Publishes reproducible hardware, model, runtime, and serving configurations
支持语言
en
已知限制
- The default comparable workload requires a 65,536-token context and a local OpenAI-compatible model server
- Results measure inference speed rather than output quality, and tool execution is skipped by default
- Managed Windows runs require an NVIDIA GPU; vLLM and SGLang recipes run only on Linux
- The public hardware and model leaderboard is still limited and is planned to expand over time






