Building an Eval Stack That Catches Regressions: Judges, RAG Metrics, and CI Gates (Part 2 of 2)
How to grade what code can't assert, evaluate RAG without debugging blind, and wire a gate that survives production.
Read moreA collection of 3 posts
How to grade what code can't assert, evaluate RAG without debugging blind, and wire a gate that survives production.
Read more
Why agent evaluation breaks the tools built for prompts, and how to measure outcomes instead of vibes.
Read more
A first-principles guide to Agent Skills: how they work, how to build one with skill-creator, and how they differ from MCP servers.
Read more