A summary note on DeepSeek’s V4 Flash 0731 release, captured from a demo post by @analogalok.
The claim
DeepSeek shipped V4 Flash 0731 on 2026-07-31, headlined by a large upgrade to agent capabilities. The eye-catching part isn’t the model itself but where it runs: quantized to 2-bit (Q2), it reportedly serves on a single consumer RTX 4090 while holding a 250k-token context.
Reported figures (single RTX 4090, Q2):
- ~12 tokens/sec generation
- 650+ tokens/sec prefill
- 250k context window
- no KV-cache quantization — the KV cache stays full precision even at that context length
Why it’s worth noting
It’s a data point on how quickly large mixture-of-experts models are collapsing onto single consumer GPUs. A 250k-context model generating on one 4090 — with an un-quantized KV cache — would have been implausible a year ago. The lever is aggressive 2-bit weight quantization holding up well enough to stay useful; that tradeoff (how far you can crush the weights before quality breaks) is the thread to follow for local and edge inference.
Caveat
These are one practitioner’s numbers from a social post, with the actual benchmarks shown in an attached video rather than a written report. Treat it as a signal about the direction of local inference, not a verified measurement.
Source: @analogalok on X, 2026-08-03