What is AssistantBench?
AssistantBench is an academic benchmark and dataset for evaluating AI web agents on realistic, time-consuming tasks such as finding gyms with specific amenities and schedules. It contains 214 tasks spanning more than 525 pages across 258 websites. The dataset and a submission leaderboard are hosted on Hugging Face, and the code, including the SeePlanAct reference agent, is on GitHub. It was created by researchers from Tel Aviv University, the Allen Institute for AI, the University of Pennsylvania, the University of Washington, and Princeton University, and published as an arXiv paper (2407.15711).
Our verdict
It suits teams that want a rigorous, multi-step web-navigation benchmark rather than single-page QA tasks, since the best reported agent at publication scored only 25.2% accuracy against much higher human performance. It carries no commercial support and exists for research evaluation, not production tool selection.
Categories:
These are research benchmarks, not products you buy. They define fixed web tasks and score how reliably an agent or model completes them, which is the closest thing the field has to an objective measure of web-agent capability. They are listed here for evaluation, not procurement.