aiagent.club
中文
GitHub repo

eval-sys/mcpmark

Python Tracked since 2025-07-01 Updated 2026-09-12 View source ↗

#756 of 803 by Stars among GitHub repos

460 Stars

MCPMark is a comprehensive, stress-testing MCP benchmark designed to evaluate model and agent capabilities in real-world MCP use.

About

An evaluation suite for agentic models in real MCP tool environments (Notion / GitHub / Filesystem / Postgres / Playwright).

MCPMark provides a reproducible, extensible benchmark for researchers and engineers: one-command tasks, isolated sandboxes, auto-resume for failures, unified metrics, and aggregated reports.

🚀 MCPMark Verified is now the default. The standard tasks in this repository are the Verified set — every environment version-pinned and every verifier stabilized. Results from earlier task versions are deprecated and not directly comparable, so please report new numbers as MCPMark Verified. On the Verified set, gpt-5.5 (xhigh) leads at 92.9% and kimi-k2.7 reaches 81.1%. See…

Excerpted from github.com/eval-sys/mcpmark

Latest metrics

Stars 460 2026-09-12
Forks 48 2026-09-12
Commits 448 2026-09-12
Releases 4 2026-09-12
Watchers 7 2026-09-12
Open issues 17 2026-09-12
Open PRs 4 2026-09-12