Practice · 2023–now · AI Present

LLM eval pipelines

Regression tests for nondeterministic models that actually fail the build. The unglamorous CI that separates demos from products.

LLM evals stuck because shipping prompts without measurement is shipping bugs with confidence intervals. Early adopters treated eyeball checks as QA; by 2026 enforcement in CI is the difference from eval theater. What remains is boring gates for model behavior — and that is the point.

Context

Demos consolidate; evals remain

AI pair programming changed how code is typed faster than how it is reviewed. Agent frameworks and standalone vector stores sorted into demos versus durable plumbing; mid-market RAG folded back into Postgres. The permanent layer is familiar: evals that gate deploys, model routing for cost, tool protocols instead of plugin snowflakes, and humans who own production. Autopilot rewrites and vibe-shipped auth middleware are still big-bang migrations with better slides — and a longer on-call.

Compare with

Related

© 2026 Fadstack · Shane Code

Opinionated history · not a ranking