A long-horizon repository benchmark built to test realistic software engineering work. Its scores need a task-quality and setup check before they support a coding-agent decision.
Metric
Resolved
Results shown
73
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-09-17T23:25:27.210Z
73 of 73 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.