Content Quality: Clear News piece (701 words, inside the 400-1200 News range) with Overview / What We Know / What We Don't Know / Analysis structure. Specific, well-attributed, appropriately hedged on the cost comparison and statistical independence.
Source Verification: Read source-0.html.gz (nature.com/articles/s41586-026-11036-y; 73k chars of extracted text) and source-1.html.gz (deepmind.google DeepNash blog; 15k chars) from disk; manifest suspicious_patterns null for both, status 200, no archive fallback. source-0 is the FULL open-access paper text (abstract, Main, Methods, ethics, author affiliations), not a stub, so all body claims are backed by the snapshot. Confirmed verbatim in source-0: published 30 September 2026; authors Sokota, Vinitsky, Hu, Fan, Kolter, Farina; affiliations CMU (Sokota, Kolter), NYU (Vinitsky), Stanford (Hu), MIT (Fan, Farina); system name Ataraxos; '20-game series, Ataraxos defeated Pim Niemeijer-the most decorated Stratego player of all time-by a margin of victory without precedent at the highest level of play: 15 wins, 1 loss and 4 draws'; 85% effective win rate counting draws as half wins; 'to our knowledge, the first superhuman result in the game's history'; series spread over 3 weeks; Pim informed Ataraxos would not adapt; payment US$1,000 + US$100/win + US$50/draw; 2025 World Championship demo 1-3 August, 40 games, 95% effective, 38 wins 2 losses 0 draws; Barrage Stratego four 50-game series all won; 'defeated PerfectDou and DouZero with statistical significance'; three-part design pattern and regularization schedule (stronger/aggressive early, weaker/smaller late); P < 2.6 x 10^-4 under the i.i.d. assumption with the caveat that outcomes are not independent; DeepNash evaluated on Gravon April 2022, 42 of 50 counted games, not top ranking; 2023 championship demo 19-9 with losses to most top players including Pim; DeepMind said DeepNash code 'is no longer functional'. Cost: 'Ataraxos RL and belief models were trained on 16 H100s for 1 week and 4 H100s for 4 days... Such a run costs less than US$8,000 at 2025 prices'; DeepNash '1,024 tensor processing unit nodes', 'to the recollection of the corresponding author of DeepNash... between 2 and 3 months... TPU v3s... Under 2025 pricing, roughly between US$3,000,000 and US$4,500,000'; 5.5 billion vs about 160 million games. The DeepNash $3-4.5M figure is therefore correctly sourced to the Nature paper (the authors' own estimate), not to the DeepMind post. source-1 (DeepMind, dated December 1, 2022): confirms DeepNash announced Dec 2022, learned Stratego from scratch to 'a human expert level', 'all-time top-three ranking among human experts' on Gravon. It states no cost, and the article does not attribute one to it.
Factual Accuracy: All numbers, names and quotes trace to the snapshots; no fabrication found. 'Most decorated' is the paper's own characterization (abstract and main text, supported there by a list of Niemeijer's titles and a quote from George Franka); the article attributes it to the paper ('described as'), which is acceptable. The headline cost is the authors' own estimate at 2025 cloud prices (paper footnote 51), not an audited or invoiced figure; the body states this clearly ('The authors write that such a run costs...', 'rests on the authors' own estimates').
Overall Assessment: Accurate, well-sourced and properly hedged report on a Nature paper whose full text was snapshotted. Headline phrasing and thin source diversity are noted but do not require a correction: the body states the cost as the authors' estimate and the headline figure matches the paper's own claim. APPROVE.