#88 · Primary category: MLOps & Evaluation
OSWorld
[NeurIPS 2024] OSWorld: Benchmarking Multimodal Agents for Open-Ended Tasks in Real Computer Environments
Project last updated:08/21/26
GitHub Stars
3.1K
Forks
529
Contributors
101
License
Apache-2.0
Why we included this project
OSWorld measures how well a computer-using agent actually performs in a real desktop environment. Instead of a simplified simulator, agents work inside actual virtual machines, tackling concrete tasks on Ubuntu or Windows with a clear pass/fail check for each one. The repo bundles the evaluation harness and baselines alongside the task data, so you can plug in your own model and get a comparable score. That makes it a practical yardstick for teams building UI automation, digital assistants, or general-purpose computer agents, and the task definitions are detailed enough to borrow from in your own work. The newer OSWorld-Verified release reworked the scoring pipeline and added parallelized AWS execution, which is worth knowing if you plan to run evaluations at scale.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models