webscrape.dev

Benchmark for evaluating computer-use AI agents on real tasks

Nathan Kessler
By Nathan KesslerUpdated

Each tool is evaluated against our methodology using public docs, vendor demos, and hands-on testing.

OSWorld website

What is OSWorld?

OSWorld is an academic benchmark and evaluation environment for testing multimodal, computer-use AI agents on real, open-ended tasks across Ubuntu, Windows, and macOS. Tasks span desktop apps, web apps, and cross-application workflows. It provides 369 tasks (361 usable), evaluated through execution-based, reproducible checker scripts, and at publication humans succeeded on over 72% of tasks versus about 12.24% for the best model. Researchers from the University of Hong Kong, Salesforce Research, Carnegie Mellon University, and the University of Waterloo built it, and the source code and task configs are public on GitHub under CC BY-SA 4.0.

Our verdict

OSWorld's execution-based checkers and multi-OS coverage give a more realistic test of computer-use agents than web-only benchmarks, though running it means standing up VM environments (VirtualBox or VMware) locally or in the cloud rather than just calling an API. The 72% versus 12.24% human-model gap comes from the original paper, so any specific score should be read against that publication point rather than as a fixed ceiling.

Categories:

These are research benchmarks, not products you buy. They define fixed web tasks and score how reliably an agent or model completes them, which is the closest thing the field has to an objective measure of web-agent capability. They are listed here for evaluation, not procurement.

Share:

Similar tools

See Benchmarks

Self-hosted, open-source benchmark for autonomous web agents

FreeJul 2026Benchmarks

Benchmark of 2,350 tasks across 137 real websites for web agents

FreeJul 2026Benchmarks

Benchmark for evaluating general AI agents on reasoning and tool use

FreeJul 2026Benchmarks

How OSWorld compares

WebArena

WebArena is the browser-only counterpart: it tests self-hosted web tasks rather than full-OS desktop environments.

Mind2Web

Mind2Web measures web-agent generalization across real websites, while OSWorld extends evaluation beyond the browser to full desktop and OS-level workflows.

GAIA

GAIA is a broader benchmark for general AI assistants that use tools such as web browsing, while OSWorld focuses specifically on operating computer GUIs across desktop, web, and file tasks.

Visit

OSWorld

Visit