If this repository is useful, please consider starring β it.
Items with π indicate open source projects.
AUTO-GENERATED FILE - DO NOT EDIT MANUALLY. Auto-generated by CI workflow or pre-commit hooks using
node generate-readme.js.
- 2026-09-22 - Corelayer (Incident Response)
- 2026-09-22 - πDarkmoon (Security)
- 2026-09-22 - πkprompt (AIOps)
Jump to: Incident Response | Observability | AIOps | IDP | IaC | Security | Deployment
| Name | Summary | Deployment | Links |
|---|---|---|---|
| AlertD | AlertD is an agentic AI teammate for SRE and DevOps on AWS, cutting alert noise and dashboard fatigue while delivering contextual answers and automated actions. | SaaS | |
| πAurora | Open source AI SRE agent that autonomously investigates incidents and delivers root cause analysis across AWS, Azure, GCP, and Kubernetes. | On-Prem | |
| AutonomOps AI | Autonomous operations platform that applies AI to improve SRE and incident management workflows. | SaaS | |
| Azure SRE Agent | AI-powered reliability assistant for Azure that automates incident response, root-cause analysis, and mitigation workflows. | SaaS | |
| Bacca.ai | AI SRE for high-scale platforms that uses tribal knowledge to triage and mitigate incidents accurately. | SaaS | |
| Beeps | AI-powered operations assistant focused on helping teams handle alerts and incident workflows faster. | SaaS | |
| Cleric | Cleric is a self-learning AI SRE that autonomously investigates production alerts and returns evidence-linked findings in Slack. It builds operational memory from prior investigations and engineer feedback while integrating with existing observability, CI/CD, and incident tooling. | SaaS | |
| Corelayer | Corelayer is an AI SRE platform for production incident investigation and prevention. It connects code, deployments, infrastructure, telemetry, and operational context so teams can investigate failures across the full production stack. | Multi | |
| DrDroid | AI that understands your production system, infrastructure, applications, and business context to investigate incidents and explain root causes. | SaaS | |
| FireHydrant | FireHydrant is an incident management platform for planning, responding to, and learning from incidents, with on-call, alerting, collaboration, and AI-assisted workflows. | SaaS | |
| Harness AI-SRE | Harness AI SRE connects alerts with deployments, feature flags, infrastructure changes, and monitoring signals to help teams investigate, document, and remediate incidents. | SaaS | |
| incident.io Investigations | Investigations is an AI capability inside incident.io's hosted incident-management platform, not a separate standalone agent. It starts from alerts and combines telemetry, code changes, past incidents, and Slack context to surface likely causes and recommended next steps. | SaaS | |
| IncidentFox | IncidentFox is an AI incident investigator that learns a team's systems and past incidents, investigates alerts, and prepares fixes for human approval. | SaaS | |
| NeuBird AI | NeuBird is an agentic reliability platform that queries telemetry in place, remembers investigation conclusions, and serves engineers and their agents through governed, policy-controlled interfaces. | SaaS | |
| NOFire AI | NOFire handles alerts, flags risky changes, turns incidents and tribal knowledge into lasting reliability memory. | SaaS | |
| PagerDuty SRE Agent | Transform critical operations with PagerDuty's AI first Operations Platform. Harness agentic AI and automation to accelerate work and build resilience. | SaaS | |
| Phoebe | The immune system for your software. AI agents that continuously investigate live data, diagnose emerging issues and generate preemptive fixes. | SaaS | |
| ProdRescue AI | Automates incident reports and evidence-backed RCA for SRE teams from Slack war rooms or logs in minutes. | SaaS | |
| Resolve AI | Resolve AI is a production-operations platform with agents for on-call triage, incident investigation, and recurring operational work. Its agents work across code, infrastructure, telemetry, and knowledge using evidence-based, multi-agent investigations. | Multi | |
| RobinRelay | RobinRelay is a Slack-native AI on-call copilot that monitors alert channels, adds context from past investigations and fixes, and answers questions about incident history. | SaaS | |
| Rootly | The all-in-one incident management platform, including AI SRE agentsβbuilt for fast-moving engineering teams to detect, manage, learn from, and resolve incidents faster. | SaaS | |
| RubixKube | RubixKube is a Site Reliability Intelligence platform for detecting anomalies, diagnosing root causes, and resolving infrastructure failures. Its platform builds persistent context from operational signals and incident investigations. | Multi | |
| RunLLM | The AI SRE for mission-critical systems that delivers transparent investigations, evidence-backed root cause analysis, and continuous runbook improvement. | SaaS | |
| Scoutflo | Scoutflo is an AI SRE platform for Kubernetes and cloud-native teams. It helps investigate alerts, identify likely root causes, and turn resolved incidents into reusable playbooks. | SaaS | |
| Sherlocks.ai | Sherlocks AI is an AI SRE teammate that investigates incidents across connected data sources and reports root cause findings with the command output and evidence behind them. | SaaS | |
| Steadwing | Steadwing is an autonomous on-call engineer that correlates evidence across a team's stack, produces actionable root cause analyses, and prepares or performs approved remediation. | SaaS | |
| TierZero AI | TierZero's AI agents investigate incidents, triage alerts, and fix production problems automatically β so your engineers can ship faster. | SaaS | |
| πTracer | OS-level AI SRE platform for high-compute workloads that accelerates alert investigation, root-cause analysis, and mitigation inside your environment. | On-Prem | |
| Traversal | Traversal is an enterprise AI SRE platform for automated alert triage, incident investigation, root-cause analysis, and remediation. Its agents combine causal machine learning with production telemetry and support read-only access and flexible on-premise deployments. | On-Prem | |
| Vibranium Labs | Vibranium Labs develops Vibe OnCall, an incident-response platform with AI agents for paging, triage, root-cause investigation, remediation, and incident knowledge capture. | SaaS | |
| Vigiles | Incident management platform for modern teams with outage detection, on-call alerting, response coordination, status pages, and AI postmortems. | SaaS | |
| Wild Moose | Wild Moose is an always-on AI first responder that gathers context from fragmented observability tools and produces evidence-backed root cause summaries and recommended next actions. | SaaS |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| Agent0 by Dash0 | Agent0 is Dash0's production AI built into its OpenTelemetry-native hosted observability platform. It surfaces production issues, investigates telemetry and code, and creates validated artifacts such as dashboards, alerts, and pull-request drafts. | SaaS | |
| Better Stack AI SRE | Better Stack AI SRE is a chat-based assistant within Better Stack's hosted observability and incident-management platform, rather than a separate standalone agent. It investigates using Better Stack telemetry and connected tools, while write actions require human approval. | SaaS | |
| Causely | Causely maintains a causal model of distributed systems so engineers and software agents can trace symptoms to likely root causes and act on remediation context through its product and MCP server. | Multi | |
| DagKnows, Inc | AI operations company focused on improving incident diagnostics and reliability workflows. | SaaS | |
| Datadog (Bits AI) | Datadog provides a cloud platform for observability and security across infrastructure, applications, logs, and user experience. Bits AI adds agents for chat, investigation, and remediation within Datadog workflows. | SaaS | |
| Deductive AI | Deductive is an AI SRE platform that investigates and resolves incidents using context from code, telemetry, engineering discussions, and knowledge bases. It builds a knowledge graph to connect relationships across those sources. | SaaS | |
| Deeptrace | Deeptrace investigates production alerts by reasoning across observability data, telemetry, and code, then helps teams move from diagnosis to fix. | SaaS | |
| Edge Delta | Observability pipeline and AI analytics platform for processing telemetry at scale and accelerating incident investigation. | SaaS | |
| Elastic | Elastic Observability brings logs, metrics, traces, and application performance data together for monitoring, investigation, and AI-assisted remediation. | SaaS | |
| Lightrun | Lightrun describes an AI SRE workflow for handling alerts and using live runtime context during development. | SaaS | |
| Logz.io | Stop Chasing Alerts. Get Ahead of Problems with AI-Powered Observability. | SaaS | |
| Metoro | Metoro is an AI SRE and Kubernetes observability platform that collects telemetry with eBPF and uses AI agents to detect production issues, investigate alerts, verify deployments, identify root causes, and generate fix PRs. | Multi | |
| Mezmo | Combine intelligent telemetry with AI-driven observability to detect issues, pinpoint root cause, and power agentic operations across logs, metrics, and traces. | SaaS | |
| Observe, Inc. | Observe is a modern observability platform built on a streaming data lake, for faster search and correlation at lower cost. | SaaS | |
| Oodle | Oodle is a unified observability platform with an AI assistant that investigates alerts across metrics, logs, traces, events, and service dependencies. It is available as managed SaaS, with customer-owned storage, or in a customer VPC. | Multi | |
| Sentry | Application performance monitoring for developers and software teams to see errors more clearly, solve issues faster, and improve reliability continuously. | SaaS | |
| SIXTA | AI-powered root cause analysis for database reliability | SaaS |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| BigPanda | AIOps platform for event correlation, incident detection, and response orchestration across modern IT operations. | SaaS | |
| Ciroos | Ciroos transforms SRE with AI-driven automation, reducing toil, detecting anomalies early, and accelerating incident investigations. | SaaS | |
| Cloudship AI | CloudShip AI turns DevOps and FinOps responsibilities into deployable agents, running them on a team's infrastructure and presenting structured operational insights. | SaaS | |
| Cokpit | Cokpit is an agentic AI platform for DevOps that provisions infrastructure, builds pipelines, investigates incidents, and documents changes across connected tools. | SaaS | |
| πHolmesGPT | HolmesGPT is a CNCF Sandbox open-source SRE agent for investigating production incidents across Kubernetes, virtual machines, cloud services, databases, and other infrastructure. It can run as a CLI, HTTP server, or Kubernetes deployment and connects to operational systems through built-in toolsets. | On-Prem | |
| Hyground | Self-hosted AI SRE agent that integrates into your infrastructure for fast incident analysis and reduced manual toil. | On-Prem | |
| πIngero | Open-source eBPF agent and MCP server for GPU causal observability, helping LLM SRE agents investigate GPU stalls and map them back to Linux kernel events and CUDA call sites. | On-Prem | |
| πK8sGPT | K8sGPT is an AI-powered tool that helps diagnose and fix Kubernetes issues with intelligent insights and automated troubleshooting. | Hybrid | |
| πKagent | Open-source Kubernetes-native framework for building and running AI agents that automate DevOps operations and troubleshooting tasks. | Hybrid | |
| KnoxOps | AI-native ops agent for SREs β gives agents production execution power with safety guardrails, human-in-the-loop review, and a built-in knowledge graph of your infrastructure. | SaaS | |
| Komodor | Komodor is an agentic operations platform for production. It provides governed workflows for incident management, troubleshooting, cost optimization, change intelligence, and software operations. | SaaS | |
| πkprompt | kprompt is an open-source AI runtime for Kubernetes that turns natural-language requests into reviewable plans and requires approval before it applies changes to a cluster. | On-Prem | |
| πKubeStellar Console | Open-source multi-cluster Kubernetes dashboard with AI-powered operations via an MCP server that bridges kubeconfig contexts to LLM agents. | Hybrid | |
| NudgeBee | Agentic AI platform for SRE & CloudOps, troubleshooting, cost optimization, and no-code workflow automation. | SaaS | |
| πObot | Open source agent platform for creating, running, and integrating autonomous assistants across workflows. | Hybrid | |
| Opsy | AI-powered reliability operations platform for faster incident response and SRE workflow automation. | SaaS | |
| Robusta Dev | Robusta provides an AI SRE platform that groups alerts, investigates incidents, identifies root causes, and recommends or executes fixes. It supports Kubernetes, cloud, and legacy environments through alert and data-source integrations. | Multi | |
| RunWhen | RunWhen is an AI SRE platform whose foreground and background agents select and run diagnostic automation inside connected environments, then return evidence-backed findings and remediation guidance. It supports hosted, hybrid, and self-hosted deployments. | Multi | |
| SRE Bench | Evaluation and benchmarking platform for SRE agents and operational AI reliability workflows. | SaaS | |
| SRE.ai | SRE.ai provides a command center and AI teammates for enterprise software delivery. Its platform covers monitoring, documentation, build guidance, testing, release orchestration, and proactive issue handling. | SaaS | |
| πStakpak | An open source agent that lives on your machines 24/7, keeps your apps running, and only pings when it needs a human. | SaaS | |
| StarSling | Multi-agent automation platform that orchestrates AI workflows for operations, troubleshooting, and remediation. | SaaS |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| StackGen | Autonomous infrastructure platform powered by Aiden for platform engineering, DevOps, and SRE teams to automate provisioning, governance, and operations. | Hybrid |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| Ops0 | ops0 is preventive cloud security for infrastructure. It finds risks in live cloud environments and routes governed, cost-aware fixes through policy, approval, pull request, and audit workflows. | SaaS |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| Cloudgeni | AI-powered cloud infrastructure platform that detects misconfigurations, remediates security and compliance issues, and generates reviewable infrastructure changes through deterministic workflows. | SaaS | |
| πDarkmoon | Darkmoon is an open-source, self-hosted platform for autonomous penetration testing across web, API, cloud, Kubernetes, Active Directory, and network targets. | On-Prem |
| Name | Summary | Deployment | Links |
|---|---|---|---|
| Cutover | Cutover's cloud-hosted Collaborative Automation platform connects teams and technology, helping you manage disaster recovery, migration, and release. | SaaS | |
| Lens K8s IDE | Kubernetes IDE for cluster operations and troubleshooting with AI-assisted diagnostics via Lens Prism. | Hybrid | |
| πSkyflo.ai | Skyflo is an open-source AI agent for DevOps and cloud operations. It plans, executes, and verifies infrastructure changes across Kubernetes, CI/CD, and cloud platforms. | Hybrid |
