aiagent.club
English
GitHub

mcpmark

观察池 · 暂无正式排名

该项目当前未满足“两个有效维度 + 两种数据源”的主榜门槛。

未排名
可观测采用度缺失
动量当前有效 · 2026-09-12
0
关注度当前有效 · 2026-09-12
0
信号可信度 依据当前有数据的独立评分维度数量计算。
2/3 · 1 种数据源

方法论 v2.0 · 快照 2026-09-12 · 超过 2 天视为过期

项目介绍

An evaluation suite for agentic models in real MCP tool environments (Notion / GitHub / Filesystem / Postgres / Playwright).

MCPMark provides a reproducible, extensible benchmark for researchers and engineers: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports.

🚀 MCPMark Verified is now the default. The standard tasks in this repository are the Verified set — every environment version-pinned and every verifier stabilized. Results from earlier task versions are deprecated and not directly comparable, so please report new numbers as MCPMark Verified. On the Verified set, gpt-5.5 (xhigh) leads at 92.9% and kimi-k2.7 reaches 81.1%. See…

各数据源

460 Star
  • Star 460
  • Fork 48
  • 提交 448
  • 发布 4