Practice · 2024–now · AI Present

Eval theater

LLM evals as slideware — metrics that look scientific and never gate a deploy. The foil to real eval pipelines in CI.

Why Eval theater became a costly fad

Eval theater failed because measurement without enforcement is theater. Real LLM evals stuck when they lived in CI with owners. The fad was buying an eval product to decorate a demo; the stuck practice is regression tests that can fail the build.

Cost of the fad

What Eval theater cost

Dashboards of vibe-check scores that never blocked a release. Green charts for leadership; prod still hallucinated.

Patterns

2022–now

AI Present

Part of the AI Present era (2022–now).

Read the AI Present era →

Compare with

Related

© 2026 Fadstack · Shane Code · Privacy

Opinionated history · not a ranking

Site updates

Occasional Fadstack notes. Confirm by email — this list stays off the book and advisory lists.

I use the address for Fadstack updates only. Privacy.