webscrape.dev

Academic benchmark testing AI agents on realistic web tasks

Nathan Kessler
By Nathan KesslerUpdated

Each tool is evaluated against our methodology using public docs, vendor demos, and hands-on testing.

AssistantBench website

What is AssistantBench?

AssistantBench is an academic benchmark and dataset for evaluating AI web agents on realistic, time-consuming tasks such as finding gyms with specific amenities and schedules. It contains 214 tasks spanning more than 525 pages across 258 websites. The dataset and a submission leaderboard are hosted on Hugging Face, and the code, including the SeePlanAct reference agent, is on GitHub. It was created by researchers from Tel Aviv University, the Allen Institute for AI, the University of Pennsylvania, the University of Washington, and Princeton University, and published as an arXiv paper (2407.15711).

Our verdict

It suits teams that want a rigorous, multi-step web-navigation benchmark rather than single-page QA tasks, since the best reported agent at publication scored only 25.2% accuracy against much higher human performance. It carries no commercial support and exists for research evaluation, not production tool selection.

Categories:

These are research benchmarks, not products you buy. They define fixed web tasks and score how reliably an agent or model completes them, which is the closest thing the field has to an objective measure of web-agent capability. They are listed here for evaluation, not procurement.

Share:

Similar tools

See Benchmarks

Self-hosted, open-source benchmark for autonomous web agents

FreeJul 2026Benchmarks

Benchmark of 2,350 tasks across 137 real websites for web agents

FreeJul 2026Benchmarks

Benchmark for evaluating general AI agents on reasoning and tool use

FreeJul 2026Benchmarks

How AssistantBench compares

GAIA

GAIA is the other benchmark AssistantBench's authors cite for comparison. It evaluates general AI assistants on multi-step reasoning and web tasks.

WebArena

WebArena is the other benchmark AssistantBench's authors compare themselves against. It tests agents on realistic tasks in a simulated web environment.

Mind2Web

Mind2Web is a similar dataset of real-world website tasks used to evaluate generalist web agents across many domains.

Visit

AssistantBench

Visit