A curated list of articles from data teams building analytics agents.
Recommended first reads · Internal builds · Browse by problem · Related reading · Contribute
- Anthropic — How a team makes self-service analytics work. Skills, metric definitions, documentation, and evaluations (Jun 2026), including how the team keeps them current as data models change.
- OpenAI — Inside an in-house data agent. The architecture behind context preparation, retrieval, permissions, and evaluations (Jan 2026).
- Ramp — Building one around an existing dbt project. SQL and Jinja examples for context tables and domain documentation (Oct 2025).
Read an article directly from the right column, or click a company for its full reading notes. Dates are publication dates; stacks reflect what the authors used at the time.
- Giving the agent context: Ramp, Gorgias, AngelList, Astronomer, Bitrise, BlaBlaCar, Cloudflare, incident.io, Meta, OpenAI.
- Testing answers and model changes: Gorgias, Anthropic, ClickHouse, Ramp, LinkedIn, Meta, Uber.
- Fixing mistakes and keeping context current: Bitrise, GitHub, Gorgias, incident.io, Meta, Replit, Astronomer.
- Reducing cost and latency: Cloudflare, Gorgias, Vercel.
- Getting people to use it: Alan, Astronomer, Bitrise, Faire, incident.io, Ramp, LinkedIn, Meta.
Product engineering, technical studies, and benchmarks that complement the internal builds.
-
Cube: Building an Agentic Analytics Harness ·
Vendor implementation. Errors, result limits, and permission-aware tools. Stack and notes. -
Omni: Benchmarking Omni’s agentic analytics harness ·
Vendor implementation. Testing quality, latency, and cost on a vendor’s analytics workload. Stack and notes. -
Cassis: A blank beats a guess ·
Vendor implementation. Building the first context from existing data assets, with an open-source bootstrap kit. Stack and notes. -
Cassis: Context engineering for analytics agents ·
Vendor implementation. How we structure tables, metrics, and business rules so an agent can find what it needs. Stack and notes. -
Lyft: From Day 1 to Production ·
Production engineering. Taking an existing analytics agent from a Streamlit prototype to an authenticated, observable internal application. Stack and notes. -
ClickHouse: The Agentic Analytics Benchmark ·
Benchmark. A reusable harness built from 201 real questions, with model-jury scoring, cost and latency measurement, and contamination checks. Code · Notes. -
Benchouse: The Analytics Agent Benchmark ·
Benchmark. Compare accuracy, completeness, restraint, and cost per question across analytics agents. Notes.
Know a team that has written up its internal build? Open an issue or pull request with the link and a sentence about what makes it useful. See the contribution guide for the format. Corrections and broken-link reports are welcome too.