#88 · Primary category: MLOps & Evaluation

OSWorld

agent artificial-intelligence benchmark cli code-generation gui language-model large-action-model llm multimodal natural-language-processing reinforcement-learning rpa vlm

[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments

Project last updated:08/21/26

GitHub Stars

3.1K

Forks

529

Contributors

101

License

Apache-2.0

Why we included this project

OSWorld measures how well a computer-using agent actually performs in a real desktop environment. Instead of a simplified simulator, agents work inside actual virtual machines, tackling concrete tasks on Ubuntu or Windows with a clear pass/fail check for each one. The repo bundles the evaluation harness and baselines alongside the task data, so you can plug in your own model and get a comparable score. That makes it a practical yardstick for teams building UI automation, digital assistants, or general-purpose computer agents, and the task definitions are detailed enough to borrow from in your own work. The newer OSWorld-Verified release reworked the scoring pipeline and added parallelized AWS execution, which is worth knowing if you plan to run evaluations at scale.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category