Practice · 2024–now · AI Present

Eval theater

LLM evals as slideware — metrics that look scientific and never gate a deploy. The foil to real eval pipelines in CI.

Eval theater failed because measurement without enforcement is theater. Real LLM evals stuck when they lived in CI with owners. The fad was buying an eval product to decorate a demo; the stuck practice is regression tests that can fail the build.

Cost of the fad

Dashboards of vibe-check scores that never blocked a release. Green charts for leadership; prod still hallucinated.

Patterns

Context

Demos consolidate; evals remain

AI pair programming changed how code is typed faster than how it is reviewed. Agent frameworks and standalone vector stores sorted into demos versus durable plumbing; mid-market RAG folded back into Postgres. The permanent layer is familiar: evals that gate deploys, model routing for cost, tool protocols instead of plugin snowflakes, and humans who own production. Autopilot rewrites and vibe-shipped auth middleware are still big-bang migrations with better slides — and a longer on-call.

Compare with

Related

© 2026 Fadstack · Shane Code

Opinionated history · not a ranking