Skip to content

Latest commit

Β 

History

154 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

Awesome AI SRE

If this repository is useful, please consider starring ⭐ it.

Tools

Items with πŸ’š indicate open source projects.

AUTO-GENERATED FILE - DO NOT EDIT MANUALLY. Auto-generated by CI workflow or pre-commit hooks using node generate-readme.js.

Recent Additions (Last 7 Days)

  • 2026-09-22 - Corelayer (Incident Response)
  • 2026-09-22 - πŸ’šDarkmoon (Security)
  • 2026-09-22 - πŸ’škprompt (AIOps)

Jump to: Incident Response | Observability | AIOps | IDP | IaC | Security | Deployment

Incident Response (32)

Name Summary Deployment Links
AlertD AlertD is an agentic AI teammate for SRE and DevOps on AWS, cutting alert noise and dashboard fatigue while delivering contextual answers and automated actions. SaaS Website LinkedIn
πŸ’šAurora Open source AI SRE agent that autonomously investigates incidents and delivers root cause analysis across AWS, Azure, GCP, and Kubernetes. On-Prem Website GitHub LinkedIn
AutonomOps AI Autonomous operations platform that applies AI to improve SRE and incident management workflows. SaaS Website LinkedIn
Azure SRE Agent AI-powered reliability assistant for Azure that automates incident response, root-cause analysis, and mitigation workflows. SaaS Website GitHub LinkedIn X
Bacca.ai AI SRE for high-scale platforms that uses tribal knowledge to triage and mitigate incidents accurately. SaaS Website LinkedIn
Beeps AI-powered operations assistant focused on helping teams handle alerts and incident workflows faster. SaaS Website LinkedIn
Cleric Cleric is a self-learning AI SRE that autonomously investigates production alerts and returns evidence-linked findings in Slack. It builds operational memory from prior investigations and engineer feedback while integrating with existing observability, CI/CD, and incident tooling. SaaS Website LinkedIn X
Corelayer Corelayer is an AI SRE platform for production incident investigation and prevention. It connects code, deployments, infrastructure, telemetry, and operational context so teams can investigate failures across the full production stack. Multi Website
DrDroid AI that understands your production system, infrastructure, applications, and business context to investigate incidents and explain root causes. SaaS Website GitHub LinkedIn X
FireHydrant FireHydrant is an incident management platform for planning, responding to, and learning from incidents, with on-call, alerting, collaboration, and AI-assisted workflows. SaaS Website LinkedIn
Harness AI-SRE Harness AI SRE connects alerts with deployments, feature flags, infrastructure changes, and monitoring signals to help teams investigate, document, and remediate incidents. SaaS Website GitHub LinkedIn X
incident.io Investigations Investigations is an AI capability inside incident.io's hosted incident-management platform, not a separate standalone agent. It starts from alerts and combines telemetry, code changes, past incidents, and Slack context to surface likely causes and recommended next steps. SaaS Website LinkedIn X
IncidentFox IncidentFox is an AI incident investigator that learns a team's systems and past incidents, investigates alerts, and prepares fixes for human approval. SaaS Website GitHub LinkedIn
NeuBird AI NeuBird is an agentic reliability platform that queries telemetry in place, remembers investigation conclusions, and serves engineers and their agents through governed, policy-controlled interfaces. SaaS Website LinkedIn
NOFire AI NOFire handles alerts, flags risky changes, turns incidents and tribal knowledge into lasting reliability memory. SaaS Website LinkedIn
PagerDuty SRE Agent Transform critical operations with PagerDuty's AI first Operations Platform. Harness agentic AI and automation to accelerate work and build resilience. SaaS Website GitHub LinkedIn X
Phoebe The immune system for your software. AI agents that continuously investigate live data, diagnose emerging issues and generate preemptive fixes. SaaS Website LinkedIn
ProdRescue AI Automates incident reports and evidence-backed RCA for SRE teams from Slack war rooms or logs in minutes. SaaS Website LinkedIn X
Resolve AI Resolve AI is a production-operations platform with agents for on-call triage, incident investigation, and recurring operational work. Its agents work across code, infrastructure, telemetry, and knowledge using evidence-based, multi-agent investigations. Multi Website LinkedIn X
RobinRelay RobinRelay is a Slack-native AI on-call copilot that monitors alert channels, adds context from past investigations and fixes, and answers questions about incident history. SaaS Website LinkedIn
Rootly The all-in-one incident management platform, including AI SRE agentsβ€”built for fast-moving engineering teams to detect, manage, learn from, and resolve incidents faster. SaaS Website LinkedIn
RubixKube RubixKube is a Site Reliability Intelligence platform for detecting anomalies, diagnosing root causes, and resolving infrastructure failures. Its platform builds persistent context from operational signals and incident investigations. Multi Website GitHub LinkedIn
RunLLM The AI SRE for mission-critical systems that delivers transparent investigations, evidence-backed root cause analysis, and continuous runbook improvement. SaaS Website LinkedIn
Scoutflo Scoutflo is an AI SRE platform for Kubernetes and cloud-native teams. It helps investigate alerts, identify likely root causes, and turn resolved incidents into reusable playbooks. SaaS Website LinkedIn
Sherlocks.ai Sherlocks AI is an AI SRE teammate that investigates incidents across connected data sources and reports root cause findings with the command output and evidence behind them. SaaS Website LinkedIn
Steadwing Steadwing is an autonomous on-call engineer that correlates evidence across a team's stack, produces actionable root cause analyses, and prepares or performs approved remediation. SaaS Website LinkedIn X Product Hunt
TierZero AI TierZero's AI agents investigate incidents, triage alerts, and fix production problems automatically β€” so your engineers can ship faster. SaaS Website LinkedIn
πŸ’šTracer OS-level AI SRE platform for high-compute workloads that accelerates alert investigation, root-cause analysis, and mitigation inside your environment. On-Prem Website GitHub
Traversal Traversal is an enterprise AI SRE platform for automated alert triage, incident investigation, root-cause analysis, and remediation. Its agents combine causal machine learning with production telemetry and support read-only access and flexible on-premise deployments. On-Prem Website LinkedIn X
Vibranium Labs Vibranium Labs develops Vibe OnCall, an incident-response platform with AI agents for paging, triage, root-cause investigation, remediation, and incident knowledge capture. SaaS Website LinkedIn
Vigiles Incident management platform for modern teams with outage detection, on-call alerting, response coordination, status pages, and AI postmortems. SaaS Website
Wild Moose Wild Moose is an always-on AI first responder that gathers context from fragmented observability tools and produces evidence-backed root cause summaries and recommended next actions. SaaS Website LinkedIn

Back to top ↑

Observability (17)

Name Summary Deployment Links
Agent0 by Dash0 Agent0 is Dash0's production AI built into its OpenTelemetry-native hosted observability platform. It surfaces production issues, investigates telemetry and code, and creates validated artifacts such as dashboards, alerts, and pull-request drafts. SaaS Website GitHub LinkedIn X
Better Stack AI SRE Better Stack AI SRE is a chat-based assistant within Better Stack's hosted observability and incident-management platform, rather than a separate standalone agent. It investigates using Better Stack telemetry and connected tools, while write actions require human approval. SaaS Website GitHub LinkedIn X
Causely Causely maintains a causal model of distributed systems so engineers and software agents can trace symptoms to likely root causes and act on remediation context through its product and MCP server. Multi Website LinkedIn
DagKnows, Inc AI operations company focused on improving incident diagnostics and reliability workflows. SaaS Website LinkedIn
Datadog (Bits AI) Datadog provides a cloud platform for observability and security across infrastructure, applications, logs, and user experience. Bits AI adds agents for chat, investigation, and remediation within Datadog workflows. SaaS Website GitHub LinkedIn X
Deductive AI Deductive is an AI SRE platform that investigates and resolves incidents using context from code, telemetry, engineering discussions, and knowledge bases. It builds a knowledge graph to connect relationships across those sources. SaaS Website LinkedIn
Deeptrace Deeptrace investigates production alerts by reasoning across observability data, telemetry, and code, then helps teams move from diagnosis to fix. SaaS Website LinkedIn
Edge Delta Observability pipeline and AI analytics platform for processing telemetry at scale and accelerating incident investigation. SaaS Website LinkedIn
Elastic Elastic Observability brings logs, metrics, traces, and application performance data together for monitoring, investigation, and AI-assisted remediation. SaaS Website GitHub LinkedIn X
Lightrun Lightrun describes an AI SRE workflow for handling alerts and using live runtime context during development. SaaS Website LinkedIn
Logz.io Stop Chasing Alerts. Get Ahead of Problems with AI-Powered Observability. SaaS Website LinkedIn
Metoro Metoro is an AI SRE and Kubernetes observability platform that collects telemetry with eBPF and uses AI agents to detect production issues, investigate alerts, verify deployments, identify root causes, and generate fix PRs. Multi Website GitHub LinkedIn X Product Hunt
Mezmo Combine intelligent telemetry with AI-driven observability to detect issues, pinpoint root cause, and power agentic operations across logs, metrics, and traces. SaaS Website LinkedIn
Observe, Inc. Observe is a modern observability platform built on a streaming data lake, for faster search and correlation at lower cost. SaaS Website LinkedIn
Oodle Oodle is a unified observability platform with an AI assistant that investigates alerts across metrics, logs, traces, events, and service dependencies. It is available as managed SaaS, with customer-owned storage, or in a customer VPC. Multi Website GitHub LinkedIn X
Sentry Application performance monitoring for developers and software teams to see errors more clearly, solve issues faster, and improve reliability continuously. SaaS Website GitHub X
SIXTA AI-powered root cause analysis for database reliability SaaS Website LinkedIn

Back to top ↑

AIOps (22)

Name Summary Deployment Links
BigPanda AIOps platform for event correlation, incident detection, and response orchestration across modern IT operations. SaaS Website LinkedIn X
Ciroos Ciroos transforms SRE with AI-driven automation, reducing toil, detecting anomalies early, and accelerating incident investigations. SaaS Website LinkedIn
Cloudship AI CloudShip AI turns DevOps and FinOps responsibilities into deployable agents, running them on a team's infrastructure and presenting structured operational insights. SaaS Website LinkedIn
Cokpit Cokpit is an agentic AI platform for DevOps that provisions infrastructure, builds pipelines, investigates incidents, and documents changes across connected tools. SaaS Website LinkedIn X
πŸ’šHolmesGPT HolmesGPT is a CNCF Sandbox open-source SRE agent for investigating production incidents across Kubernetes, virtual machines, cloud services, databases, and other infrastructure. It can run as a CLI, HTTP server, or Kubernetes deployment and connects to operational systems through built-in toolsets. On-Prem Website GitHub LinkedIn X
Hyground Self-hosted AI SRE agent that integrates into your infrastructure for fast incident analysis and reduced manual toil. On-Prem Website LinkedIn
πŸ’šIngero Open-source eBPF agent and MCP server for GPU causal observability, helping LLM SRE agents investigate GPU stalls and map them back to Linux kernel events and CUDA call sites. On-Prem Website GitHub X
πŸ’šK8sGPT K8sGPT is an AI-powered tool that helps diagnose and fix Kubernetes issues with intelligent insights and automated troubleshooting. Hybrid Website GitHub X
πŸ’šKagent Open-source Kubernetes-native framework for building and running AI agents that automate DevOps operations and troubleshooting tasks. Hybrid Website GitHub
KnoxOps AI-native ops agent for SREs β€” gives agents production execution power with safety guardrails, human-in-the-loop review, and a built-in knowledge graph of your infrastructure. SaaS Website X
Komodor Komodor is an agentic operations platform for production. It provides governed workflows for incident management, troubleshooting, cost optimization, change intelligence, and software operations. SaaS Website LinkedIn
πŸ’škprompt kprompt is an open-source AI runtime for Kubernetes that turns natural-language requests into reviewable plans and requires approval before it applies changes to a cluster. On-Prem Website GitHub
πŸ’šKubeStellar Console Open-source multi-cluster Kubernetes dashboard with AI-powered operations via an MCP server that bridges kubeconfig contexts to LLM agents. Hybrid Website GitHub X
NudgeBee Agentic AI platform for SRE & CloudOps, troubleshooting, cost optimization, and no-code workflow automation. SaaS Website LinkedIn
πŸ’šObot Open source agent platform for creating, running, and integrating autonomous assistants across workflows. Hybrid Website GitHub
Opsy AI-powered reliability operations platform for faster incident response and SRE workflow automation. SaaS Website GitHub
Robusta Dev Robusta provides an AI SRE platform that groups alerts, investigates incidents, identifies root causes, and recommends or executes fixes. It supports Kubernetes, cloud, and legacy environments through alert and data-source integrations. Multi Website GitHub LinkedIn X
RunWhen RunWhen is an AI SRE platform whose foreground and background agents select and run diagnostic automation inside connected environments, then return evidence-backed findings and remediation guidance. It supports hosted, hybrid, and self-hosted deployments. Multi Website GitHub LinkedIn
SRE Bench Evaluation and benchmarking platform for SRE agents and operational AI reliability workflows. SaaS Website LinkedIn
SRE.ai SRE.ai provides a command center and AI teammates for enterprise software delivery. Its platform covers monitoring, documentation, build guidance, testing, release orchestration, and proactive issue handling. SaaS Website LinkedIn
πŸ’šStakpak An open source agent that lives on your machines 24/7, keeps your apps running, and only pings when it needs a human. SaaS Website GitHub LinkedIn X
StarSling Multi-agent automation platform that orchestrates AI workflows for operations, troubleshooting, and remediation. SaaS Website LinkedIn X

Back to top ↑

IDP (1)

Name Summary Deployment Links
StackGen Autonomous infrastructure platform powered by Aiden for platform engineering, DevOps, and SRE teams to automate provisioning, governance, and operations. Hybrid Website LinkedIn

Back to top ↑

IaC (1)

Name Summary Deployment Links
Ops0 ops0 is preventive cloud security for infrastructure. It finds risks in live cloud environments and routes governed, cost-aware fixes through policy, approval, pull request, and audit workflows. SaaS Website LinkedIn X

Back to top ↑

Security (2)

Name Summary Deployment Links
Cloudgeni AI-powered cloud infrastructure platform that detects misconfigurations, remediates security and compliance issues, and generates reviewable infrastructure changes through deterministic workflows. SaaS Website LinkedIn
πŸ’šDarkmoon Darkmoon is an open-source, self-hosted platform for autonomous penetration testing across web, API, cloud, Kubernetes, Active Directory, and network targets. On-Prem Website GitHub

Back to top ↑

Deployment (3)

Name Summary Deployment Links
Cutover Cutover's cloud-hosted Collaborative Automation platform connects teams and technology, helping you manage disaster recovery, migration, and release. SaaS Website LinkedIn X
Lens K8s IDE Kubernetes IDE for cluster operations and troubleshooting with AI-assisted diagnostics via Lens Prism. Hybrid Website GitHub LinkedIn X
πŸ’šSkyflo.ai Skyflo is an open-source AI agent for DevOps and cloud operations. It plans, executes, and verifies infrastructure changes across Kubernetes, CI/CD, and cloud platforms. Hybrid Website GitHub X

Back to top ↑