GitHub
mcpmark
观察池 · 暂无正式排名
该项目当前未满足“两个有效维度 + 两种数据源”的主榜门槛。
— 未排名
可观测采用度缺失 —
动量当前有效 · 2026-09-12 0
关注度当前有效 · 2026-09-12 0
信号可信度 依据当前有数据的独立评分维度数量计算。
中2/3 · 1 种数据源
项目介绍
An evaluation suite for agentic models in real MCP tool environments (Notion / GitHub / Filesystem / Postgres / Playwright).
MCPMark provides a reproducible, extensible benchmark for researchers and engineers: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports.
🚀 MCPMark Verified is now the default. The standard tasks in this repository are the Verified set — every environment version-pinned and every verifier stabilized. Results from earlier task versions are deprecated and not directly comparable, so please report new numbers as MCPMark Verified. On the Verified set, gpt-5.5 (xhigh) leads at 92.9% and kimi-k2.7 reaches 81.1%. See…
各数据源
460 Star
- Star 460
- Fork 48
- 提交 448
- 发布 4