What is OSWorld?
OSWorld is an academic benchmark and evaluation environment for testing multimodal, computer-use AI agents on real, open-ended tasks across Ubuntu, Windows, and macOS. Tasks span desktop apps, web apps, and cross-application workflows. It provides 369 tasks (361 usable), evaluated through execution-based, reproducible checker scripts, and at publication humans succeeded on over 72% of tasks versus about 12.24% for the best model. Researchers from the University of Hong Kong, Salesforce Research, Carnegie Mellon University, and the University of Waterloo built it, and the source code and task configs are public on GitHub under CC BY-SA 4.0.
Our verdict
OSWorld's execution-based checkers and multi-OS coverage give a more realistic test of computer-use agents than web-only benchmarks, though running it means standing up VM environments (VirtualBox or VMware) locally or in the cloud rather than just calling an API. The 72% versus 12.24% human-model gap comes from the original paper, so any specific score should be read against that publication point rather than as a fixed ceiling.
Categories:
These are research benchmarks, not products you buy. They define fixed web tasks and score how reliably an agent or model completes them, which is the closest thing the field has to an objective measure of web-agent capability. They are listed here for evaluation, not procurement.