Case Study — 03
Local AI Lab

Situation
Running AI in a product usually means sending user data to an external API, paying per request and depending on a key. Open-source models have become small enough to question that default, but what actually runs well inside a browser tab is still mostly guesswork.
Task
Find out what can run entirely on the user's device, and prove it with working demos instead of benchmarks — each one paired with a short write-up of what I learned.
Action
- Built 7 demos across language, vision, audio and retrieval: sentiment analysis, zero-shot classification, translation, semantic search, image classification, object detection and speech recognition
- Ran every model inside a dedicated Web Worker so inference never blocks the interface
- Used WebGPU when available, with a transparent WASM fallback for devices that don't support it
- Loaded quantized models only on demand, showing the real download size before a single byte is fetched, and cached them in the browser
- Configured cross-origin isolation with COEP credentialless, enabling multi-threaded WASM while still loading weights from the Hugging Face CDN
- Built a shared demo kit, so each new experiment inherits the download gate, cache restore and model controls without extra work
Result
- 7 AI demos in production with no backend, no API keys and no data leaving the user's device
- Models from 24 MB to 119 MB that download once and are reused from the browser cache on every visit
- An architecture where adding a new model is a content entry and a single component, not a rewrite
Technologies
- Astro
- React
- TypeScript
- Transformers.js
- Hugging Face
- WebGPU
- Web Workers
- Vitest