DeepSeek V4 Flash 0731

@navalm 2026-08-05 tags: aillminferencedeepseek views: —

A summary note on DeepSeek’s V4 Flash 0731 release, captured from a demo post by @analogalok.

The claim

DeepSeek shipped V4 Flash 0731 on 2026-07-31, headlined by a large upgrade to agent capabilities. The eye-catching part isn’t the model itself but where it runs: quantized to 2-bit (Q2), it reportedly serves on a single consumer RTX 4090 while holding a 250k-token context.

Reported figures (single RTX 4090, Q2):

  • ~12 tokens/sec generation
  • 650+ tokens/sec prefill
  • 250k context window
  • no KV-cache quantization — the KV cache stays full precision even at that context length

Why it’s worth noting

It’s a data point on how quickly large mixture-of-experts models are collapsing onto single consumer GPUs. A 250k-context model generating on one 4090 — with an un-quantized KV cache — would have been implausible a year ago. The lever is aggressive 2-bit weight quantization holding up well enough to stay useful; that tradeoff (how far you can crush the weights before quality breaks) is the thread to follow for local and edge inference.

Caveat

These are one practitioner’s numbers from a social post, with the actual benchmarks shown in an attached video rather than a written report. Treat it as a signal about the direction of local inference, not a verified measurement.

Source: @analogalok on X, 2026-08-03