Coding Horizon

DeepSeek V4.1 Flash Makes Claude Earn Your Money

Every figure the picture puts on screen, chased to a primary source. Checked 2026-09-11.

DeepSeek V4.1 Flash benchmark table

Source: DeepSeek V4.1 Flash technical report, Table 3, and the model card benchmark table.

All figures are publisher reported by DeepSeek, at maximum reasoning effort, in DeepSeek’s own agent configuration.

Benchmark V4.1 Flash V4 Flash Claude Opus 5 GPT 5.6 Sol Kimi K3 GLM 5.3
DeepSWE v1.1 (resolved) 74.2 54.4 74.0 73.0 67.5 66.9
Terminal Bench 2.1 (pass@1) 90.6 82.7 89.1 88.8 88.3 88.2
Terminal Bench 4.0 (pass@1) 31.2 7.0 51.8 39.9 12.6 37.9
AutomationBench (pass@1) 54.8 37.7 50.3 45.8 46.7 48.8

Derived figures shown on screen, computed from the table:

The small DeepSWE difference against Opus 5 does not establish a statistically significant lead; the script describes it as neck and neck.

Architecture

Source: model card and technical report abstract and architecture sections (links above).

Arithmetic shown on screen: 890 bytes times 1,000,000 tokens = 890,000,000 bytes, about 890 MB (decimal). This is the global attention cache only, not weights, other caches or runtime memory. The V4 Flash per token figure is not printed on screen; the picture draws the old block at about four times the height, as the model card states.

Reasoning effort

Source: technical report, section 5.3 (effort control), printed page 34.

The chart in the effort beat is qualitative (accuracy rising and flattening, tokens rising steeply) with the 60 to 80 band shaded; no y axis values are printed.

API pricing

Source: DeepSeek API pricing page, https://api-docs.deepseek.com/quick_start/pricing

The $10 versus $5 example is arithmetic on identical billable usage at half rate, not a measured agent run.

License and weights

Source: model repository, https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash

Reference inference configuration

Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/main/inference/README.md

Kimi K3

Source: https://huggingface.co/moonshotai/Kimi-K3

GLM 5.3

Source: https://huggingface.co/zai-org/GLM-5.3

Qwen 3.8 27B

Source: https://huggingface.co/Qwen/Qwen3.8-27B

Arithmetic shown on screen: 27 billion parameters at 4 bits each = 27e9 x 4 / 8 = 13.5e9 bytes, about 13.5 GB, before vision components, quantisation metadata, caches and runtime overhead.

Not checked