A broad research-agent benchmark for open-ended information gathering, synthesis, and answer construction across wide search spaces.
Metric
Score
Results shown
16
Publisher
BenchLM exact-source leaderboard
Snapshot
Fetched 2026-09-17T23:25:31.424Z
16 of 16 models
Higher is better
Scores are published snapshots and are only comparable within the same benchmark, version, metric, and evaluation configuration.