#57 · Primary category: AI DevOps & SRE Tools

ongrid

ai-agents aiops alerting chatops devops golang grafana incident-response kubernetes llm llm-agent mcp monitoring observability opentelemetry prometheus rag root-cause-analysis self-hosted sre

An ops AI Agent that understands your infrastructure, finds the root cause, and fixes it — right from Slack, Telegram, Lark or DingTalk.

Project last updated:08/28/26

GitHub Stars

931

Forks

201

Contributors

18

License

AGPL-3.0

Why we included this project

Ongrid takes a different route to incident response: instead of paging a human with another alert, it starts digging on its own. Once it is wired into your Prometheus, Grafana, OpenTelemetry, and Kubernetes setup, an alert turns into an automated sweep across metrics, logs, traces, and topology, with a likely root cause surfaced alongside relevant runbooks pulled from a RAG knowledge base. It can also search code and run fixes remotely, and specialist agents and skills cover distinct domains. You drive it from Slack, Telegram, Lark, or DingTalk, and because it is written in Go and self-hosted, an SRE or platform team can keep the whole thing inside its own infrastructure instead of depending on a hosted vendor. That is what separates it from a thin chat wrapper around your dashboards.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category