Practice · 2023–now · AI Present

Local inference

Run the model on your laptop or rack. Privacy, cost control, and the refusal to send every token to a vendor.

Why Local inference stuck

Local inference stuck for privacy-sensitive and offline workflows as open weights and IDE-local models matured. It failed as a default for teams that needed frontier quality without GPU ops. The fad was "replace the API tomorrow"; the stuck practice is hybrid — local for routine, cloud for peak.

Case studies

2022–now

AI Present

Part of the AI Present era (2022–now).

Read the AI Present era →

Compare with

Related

© 2026 Fadstack · Shane Code · Privacy

Opinionated history · not a ranking

Site updates

Occasional Fadstack notes. Confirm by email — this list stays off the book and advisory lists.

I use the address for Fadstack updates only. Privacy.