Content Quality: Well-structured News piece with a clear Overview / What We Know / What We Don't Know format. Technical claims (MoE backbone size, Engram/N-gram conditional-memory module, Causal Encoder-Decoder architecture, KV-cache bytes-per-token, GPU memory math, pricing tiers, benchmark scores) are dense but each is individually attributed to a named outlet or DeepSeek's own model card, and the piece is transparent about which figures are DeepSeek's own self-reported benchmarks versus independently verified facts.
Source Verification: Read all 4 source snapshots in full (gunzip'd from sources/2026-09/deepseek-ships-v41-flash-a-763-billion-parameter-model-that-cuts-kv-cache-memory-to-a-quarter-of-its-predecessors/), verified sha256 of each decompressed file against manifest.json (all 4 matched), and confirmed all 4 fetched with status_code 200 and suspicious_patterns: null. (1) source-0.html.gz (VentureBeat) — confirmed the 552B backbone figure, 196B Engram module, 8B/16B prefill/decode activation split, 890 bytes/token KV cache (~1/4 of V4-Flash), full off-peak/peak pricing table ($0.003/$0.15/$0.60 off-peak, $0.006/$0.30/$1.20 peak) and peak-window hours (Mon-Fri 01:00-04:00 and 06:00-10:00 UTC), rival pricing comparisons (GPT-5.6 Sol, Claude Opus 5, Kimi K3), MIT license claim, the DeepSWE/CyberGym/AutomationBench/Terminal-Bench/GPQA/SEC-Bench benchmark figures, the 284B/13B V4-Flash predecessor figures, the OpenDesign third-party quote ('a narrow third-party workload, not a general model evaluation') verbatim, the robustness-limitation quote ('could cause capability degradation in untested edge cases') verbatim, and the Hacker News commentary paraphrases. All confirmed accurate as attributed. (2) source-1.html.gz (The Register) — confirmed the 763-billion-parameter total figure ('At 763 billion parameters, the point release is more than 2.5x the size of the model it replaces'), the 196B N-gram/Engram figure, the 13%-25% KV-cache-of-predecessor figure and 'four to eight times as many users' quote verbatim, the 763GB/567GB FP8 GPU-memory figures verbatim, the Gemma Per-Layer-Embedding comparison, and the Alibaba Qwen 3.8-Flash-Next (180B backbone, 51B N-gram pool) industry-parallel paragraph. All confirmed accurate as attributed. (3) source-2.html.gz (Hugging Face model card) — this is the primary technical source and independently corroborates the 552B backbone, 196B Engram module, 40-layer (20+20) Causal Encoder-Decoder split, 8B/16B activation split, and 890-bytes/token KV cache figures cited from VentureBeat. The page's own 'Safetensors Model size' field reads '763B params', directly corroborating the total-parameter figure used in the headline (552B backbone + 196B Engram + ~15B in vision encoder/projector/embeddings not broken out elsewhere = 763B total, per DeepSeek's own auto-computed safetensors size) rather than being a bot fabrication. The full evaluation tables here independently confirm every benchmark number in the article's Benchmarks bullet (DeepSWE v1.1: Opus-5.0 74.0, GPT-5.6 Sol 73.0, DS-V4.1-Flash 74.2; CyberGym DS-V4.1-Flash 88.1; AutomationBench DS-V4.1-Flash 54.8; Terminal-Bench 3.0 Opus-5.0 43.3 vs DS-V4.1-Flash 30.0; Terminal-Bench 4.0 Opus-5.0 51.8 vs DS-V4.1-Flash 31.2; GPQA Diamond and SEC-Bench Pro both led by GPT-5.6 Sol) and the V4-Flash-Base predecessor row (284B backbone / 13B activated). (4) source-3.html.gz (DeepSeek API pricing docs) — confirmed the exact off-peak/peak pricing figures used in the Pricing bullet verbatim against the live 'Models & Pricing' table. However, this same page directly contradicts a separate claim in the article's Migration bullet: the article states, per VentureBeat's Sept. 10 report, that 'starting September 14, calls to deepseek-v4-pro will also route to V4.1-Flash until a future V4.1-Pro model arrives.' The pricing page fetched for this review (Sept. 16, after the Sept. 14 date) states in footnote (2) on the deepseek-v4-pro row: 'In response to user demand, we have decided to continue providing API services for DeepSeek V4 Pro after September 14, 2026, with the billing method remaining unchanged.' DeepSeek reversed the announced V4-Pro retirement between VentureBeat's Sept. 10 report and the Sept. 16 submission date. This is a genuine factual issue, not a bot fabrication — the claim was accurately attributed to VentureBeat's reporting at the time, but is now superseded by DeepSeek's own current documentation (which the bot cited as source #4 for pricing but apparently did not cross-check for the migration claim).
Factual Accuracy: Every specific in the article — the 763B total parameter count, 552B backbone, 196B Engram module, 890 bytes/token KV cache, 763GB/567GB GPU memory figures, all pricing tiers, all benchmark scores, the 284B/13B predecessor figures, and all direct quotes inside quote marks — was independently verified against the underlying snapshots and traces to a cited source. No hallucinated quotes or misattributions found. One subordinate claim (the deepseek-v4-pro migration date/plan) is accurately attributed to VentureBeat's reporting but has since been superseded by DeepSeek's own current pricing documentation, which the article does not reflect. This is the one issue driving the APPROVE_WITH_CORRECTIONS verdict.
Overall Assessment: A well-sourced, well-organized News piece on a legitimate, non-duplicate release. All headline, summary, and lead figures (763B parameters, KV-cache reduction, pricing) are independently corroborated across multiple sources including DeepSeek's own model card and pricing docs. The single issue found is a subordinate claim in the body (the deepseek-v4-pro migration timeline) that was accurately reported at the time by VentureBeat but has since been superseded by DeepSeek's own current documentation. This is exactly the kind of single, honestly-correctable subordinate issue APPROVE_WITH_CORRECTIONS is designed for — it does not touch the headline, summary, or lead, and a single corrections note fully informs readers. Verdict: APPROVE_WITH_CORRECTIONS.