evals
Here are 1,327 public repositories matching this topic...
AI Observability & Evaluation
-
Updated
Aug 13, 2026 - Python
Python SDK for AI agent monitoring, LLM cost tracking, benchmarking, and more. Integrates with most LLMs and agent frameworks including CrewAI, Agno, OpenAI Agents SDK, Langchain, Autogen, AG2, and CamelAI
-
Updated
Jun 25, 2026 - Python
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
-
Updated
Aug 12, 2026 - Python
AI observability platform for production LLM and agent systems.
-
Updated
Aug 12, 2026 - Python
Framework for evaluating and improving agents
-
Updated
Aug 13, 2026 - Python
Evaluation and Tracking for LLM Experiments and AI Agents
-
Updated
Aug 12, 2026 - Python
Laminar - open-source observability platform purpose-built for AI agents. YC S24.
-
Updated
Aug 12, 2026 - TypeScript
AI system design guide for engineers building production AI systems and evals.
-
Updated
Jul 31, 2026
Testing and evaluation platform to chat, inspect, and debug MCP servers, MCP apps, and ChatGPT apps.
-
Updated
Aug 13, 2026 - TypeScript
OpenSource Production ready Customer service with built in Evals and monitoring
-
Updated
Jun 18, 2026 - TypeScript
Waku Waku! Waku Agent is a local-first AI agent harness you actually own, including loop, memory, eval, all in code built to stay legible as it grows.
-
Updated
Aug 12, 2026 - Python
🥤 RAGLite is a Python toolkit for Retrieval-Augmented Generation (RAG) with DuckDB or PostgreSQL
-
Updated
Jul 9, 2026 - Python
Build agentic systems. Run them with confidence. Orchestrate agents, automate business processes, inspect every execution, and keep humans in control. Deploy Heym on your own infrastructure.
-
Updated
Aug 12, 2026 - Python
A skill creator that proves its skills work. Evidence-driven skill creation for Claude Code and Codex: baseline-tested generation, per-skill regression evals, ecosystem doctor, cross-runtime compile, and an opt-in proactive advisor.
-
Updated
Jul 29, 2026 - Python
A curated, non-BS library of the best resources for building and evaluating AI agents — papers, blogs, talks, tools, benchmarks. Maintained by BenchFlow.
-
Updated
Jul 31, 2026
Trace-native CI/CD for AI agents — production failures become regression tests that block the PR. Auto-detect, cluster, freeze into hermetic cases, replay in CI for $0.
-
Updated
Aug 12, 2026 - Python
Agent trace and tool-use safety evaluation lab.
-
Updated
May 2, 2026 - Python
Improve this page
Add a description, image, and links to the evals topic page so that developers can more easily learn about it.
Add this topic to your repo
To associate your repository with the evals topic, visit your repo's landing page and select "manage topics."