Real-world multi-step tool-use evaluation through the Model Context Protocol.
Metric
Pass rate
Results shown
35
Publisher
Scale Labs leaderboard
Snapshot
Fetched 2026-09-17T23:25:26.942Z
35 of 35 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.