AI & Machine Learning
244 articles RSS
TrueFoundry Open-Sources TrueForge, an Agent Harness Benchmarked Up to 75% Cheaper Than Claude Managed Agents
TrueFoundry's MIT-licensed TrueForge harness matched Claude Managed Agents on a 14-task enterprise benchmark while costing up to 75% less per run.
OpenAI Reinstates Five-Hour Usage Limit on Codex and ChatGPT Work for Plus Subscribers
OpenAI is bringing back a five-hour usage cap on Codex and ChatGPT Work for Plus subscribers starting August 25, citing compute load management.
LinkedIn's Multi-Agent AI Code Review System Posts 79,000 Reviews a Week, With Developers Accepting 64% of Suggestions
LinkedIn detailed a Kubernetes-based multi-agent AI code review platform that completes 79,000+ reviews weekly, with a 63.9% developer acceptance rate.
Study Finds AI Code Review Bots Already Grading Other AI Agents' GitHub Pull Requests at Scale
A new empirical study finds 248,641 AI-authored GitHub pull requests have already received an AI-generated review, with distinct patterns by agent pairing.
Study of 33,097 Agentic Pull Requests Finds AI Coding Agents Favor AGENTS.md Over README and API Docs
A new empirical study finds coding agents overwhelmingly read and write their own instruction files, rarely touch classical documentation, and almost never use docs to recover from failures.
Cloudflare Brings WriteGuard MCP Server Controls Out of Internal Rollout Into Private Beta
Cloudflare opens a private beta of WriteGuard, a policy and audit layer that classifies and can block risky write actions AI agents take through MCP servers.
AMD Ships GAIA 0.23, Adding Terminal-Native Agent Installs and MCP Security Hardening for Ryzen AI NPUs
AMD's open-source local AI agent framework for Ryzen AI NPUs gains a terminal-native agent hub, cross-surface confirmation gates, and MCP security fixes.
Researchers Show Coding-Agent Task Difficulty Is Predictable From Static Code Structure Alone, Hitting 0.863 AUC
A George Mason University study finds a software task's difficulty for AI coding agents can be predicted from patch and repository structure before the agent ever runs.
Google Researchers Show Spec-Driven Test Generation Cuts Missed Bugs in AI Coding Agents
A Google study finds AI agents that first write a semi-formal specification before generating tests catch 9.8 percentage points more real bugs than agents that generate tests directly.
Researchers Coin 'Coherence Debt' to Explain Why Coding Agents Fabricate Fixes Instead of Asking When Facts Go Missing
A study spanning seven AI models and five coding-agent harnesses finds they fail identically when a required repository fact is unavailable, and fabricate rather than stop.
Alibaba Releases Qwen3.8-27B Open Weights, a Dense Coding Model That Beats Claude Opus 4.6 Max on Two Benchmarks
Alibaba's Qwen team shipped Qwen3.8-27B as Apache 2.0 open weights, the on-premise checkpoint promised alongside Qwen3.8-Max, with coding benchmarks that top Claude Opus 4.6 Max on two of four tests.
New Study Finds AI-Generated Bug-Fix Patches Far More Bloated Than Developers', Proposes RECAP to Slim Them Down
A new study finds AI-written bug-fix patches are far larger and more complex than developers' own fixes even when tests pass, and proposes a lightweight adapter called RECAP to trim the bloat.