#61 · Primary category: AI Agents & Automation

UI-TARS-desktop

agent agent-tars browser-use computer-use cowork gui-agent gui-operator mcp mcp-server multimodal tars ui-tars vision vlm

The Open-Source Multimodal AI Agent Stack: Connecting Cutting-Edge AI Models and Agent Infra

Project last updated:08/05/26

GitHub Stars

38.7K

Forks

3.9K

Contributors

49

License

Apache-2.0

Why we included this project

Most agent frameworks stop at text; this one is built to see and click. The repo bundles a desktop app that turns the UI-TARS vision-language model into a local computer and browser operator, plus a broader agent stack with a CLI and web UI for driving tasks across your terminal, browser, and product. It handles real GUI work like clicking through settings, filling forms, and navigating pages by combining visual grounding with DOM-based control, and it connects to MCP servers so the agent can reach external tools. Teams evaluating computer-use automation get a cross-platform starting point that runs fully locally for privacy-sensitive workflows, with the option to point it at remote models when you need more horsepower.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category