#57 · Primary category: AI DevOps & SRE Tools
ongrid
An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.
Project last updated:08/28/26
GitHub Stars
931
Forks
201
Contributors
18
License
AGPL-3.0
Why we included this project
Ongrid takes a different route to incident response: instead of paging a human with another alert, it starts digging on its own. Once it is wired into your Prometheus, Grafana, OpenTelemetry, and Kubernetes setup, an alert turns into an automated sweep across metrics, logs, traces, and topology, with a likely root cause surfaced alongside relevant runbooks pulled from a RAG knowledge base. It can also search code and run fixes remotely, and specialist agents and skills cover distinct domains. You drive it from Slack, Telegram, Lark, or DingTalk, and because it is written in Go and self-hosted, an SRE or platform team can keep the whole thing inside its own infrastructure instead of depending on a hosted vendor. That is what separates it from a thin chat wrapper around your dashboards.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
harness
Harness Open Source is an end-to-end developer platform with Source Control Management, CI/CD Pipelines, Hosted Developer Environments, and Artifact Registries.
nx
The Monorepo Platform that amplifies both developers and AI agents. Nx optimizes your builds, scales your CI, and fixes failed PRs automatically. Ship in half the time.
jenkins
Jenkins automation server
awx
AWX provides a web-based user interface, REST API, and task engine built on top of Ansible. It is one of the upstream projects for Red Hat Ansible Automation Platform.
dagger
Automation engine to build, test and ship any codebase. Runs locally, in CI, or directly in the cloud