🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
-
Updated
Oct 5, 2026 - TypeScript
🪢 Open source agent evals & observability: Trace, evaluate, and improve LLM applications with one open platform.
Debug, evaluate, and monitor your LLM applications, RAG systems, and agentic workflows with comprehensive tracing, automated evaluations, and production-ready dashboards.
AI Agent Engineering Platform built on an Open Source TypeScript AI Agent Framework
Fastest enterprise AI gateway (50x faster than LiteLLM) with adaptive load balancer, cluster mode, guardrails, 1000+ models support & <100 µs overhead at 5k RPS.
Connect Your Agents And Harnesses With Any Provider 🦚
🧊 Open source LLM observability platform. One line of code to monitor, evaluate, and experiment. YC W23 🍓
Next-generation AI Agent Optimization Platform: Cozeloop addresses challenges in AI agent development by providing full-lifecycle management capabilities from development, debugging, and evaluation to monitoring.
Observability and enforcement for AI agent harnesses. Capture every run and runtime reliability with policy enforcement.
Open-source observability for AI agents. Find where your agents fail, dispatch your coding agent to fix it, and verify the fix against real traces.
AI observability platform for production LLM and agent systems.
Agent Skills as a Memory Layer
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
OpenLIT is the open-source agent harness engineering platform: trace, evaluate, guard, and improve everything around the model in your AI agents, on OTEL.
Open-source, end-to-end platform for evaluating, observing, and improving LLM and AI agent applications. Tracing · Evals · Simulations · Datasets · Gateway · Guardrails. Self-hostable. Apache 2.0.
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
The Continuous-Improvement Stack for Agents. Our environment data and evals power agent improvement and monitoring.
Local gateway for Claude Code, Codex and other AI clients on macOS, Windows and Linux: switch upstreams without touching clients, keep API keys from relays, cut off malicious tool calls, trace every request.
Self-hosted AI API and MCP gateway for organizations: SSO and RBAC, per-user identity for MCP tool calls, PII redaction and tool-call guards, rate limits, budgets, cost accounting and audit logs.
Local-first observability for AI coding agents. Every Claude Code and Codex session's tokens, cache, models and cost, on this machine and every machine you connect through a self-hosted team hub. Presenting mode, policy hooks, Apple-signed and notarized macOS builds, attested releases. MIT.
TraceRoot - open source self improving layer for ai agents YC S25
To associate your repository with the llm-observability topic, visit your repo's landing page and select "manage topics."