#34 · Primary category: AI DevOps & SRE Tools

holmesgpt

aiops chatbot chatops devops devops-tools incident incident-management incident-response jira kubernetes llm llm-agent llm-framework llms monitoring observability prometheus site-reliability-engineering slack sre

SRE Agent - CNCF Sandbox Project

Project last updated:08/26/26

GitHub Stars

3.2K

Forks

457

Contributors

74

License

Apache-2.0

Why we included this project

Investigating a production outage usually means stitching together dashboards, logs, and alerts by hand. HolmesGPT automates that digging: it connects to your observability stack, Kubernetes, cloud providers, and databases, pulls the relevant metrics and logs, and works toward a root cause instead of just surfacing a status page. Its Operator mode runs continuously in the background, spots problems before customers do, and can message Slack or even open a pull request to apply a fix. It also writes findings back to Jira, PagerDuty, or OpsGenie, so the investigation results fit into your existing on-call flow. For SRE and platform teams that want an LLM agent operating inside real infrastructure rather than in a sandbox, it's worth trying against your own incident history.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category