
How we self-host DeepSeek V4 Flash on AWS spot instances
We recently launched a free tier powered by a self-hosted DeepSeek V4 Flash. Here's how we picked the model and hardware together, adapted an open source vLLM recipe, and made the KV cache — not raw compute — the thing that let the economics work.
Engineering














