Agents
65 articles RSS
Cognition Crosses $1 Billion in Annualized Revenue, Delivering on a Target Set at Its May Funding Round
The Devin AI-coding startup hit $1B in annualized revenue run rate, weeks after a $2B Series E valued it at $48B.
OpenAI Launches Agents API in Public Beta, Opening the Codex Agent Harness to Developers
OpenAI's new Agents API, in public beta, turns the managed harness behind Codex into infrastructure any developer can build custom AI agents on, drawing mixed analyst reaction.
Claude Becomes the Fifth Major AI Chatbot App to Gain Apple CarPlay Support
Anthropic added CarPlay integration to its Claude iOS app, letting drivers hold hands-free voice conversations, following ChatGPT, Perplexity, Grok, and Meta AI.
Anthropic Previews a Model Hardware Standard Letting AI Agents Directly Operate Lab Robots and Manufacturing Equipment
Anthropic opened a research preview of the Model Hardware Standard, letting AI agents read and write to lab and factory hardware through standardized drivers, cutting integration from weeks to hours in early tests.
MCR-Bench Study Finds Leading LLMs' Code-Review Accuracy Collapses as Review Rounds Pile Up
A new benchmark testing seven LLMs on real multi-round GitHub code reviews finds accuracy drops sharply as review rounds increase, exposing weak memory across rounds.
AWS Open-Sources Kiro Crew for Asynchronous AI Coding Agents
AWS open-sourced Kiro Crew, an agent-orchestration tool already used internally at Amazon by more than 39,000 developers, under an Apache 2.0 license.
Meta and UIUC Researchers Get an 8-Billion-Parameter Model to Match Claude Opus 4.5 With a Smarter Agent Harness
EvoHarness-RL trains Qwen3-8B to hit 96.9% on ALFWorld, edging out Claude Opus 4.5's 96.4% baseline by rethinking agent memory and state.
New SWE-Prime Method Trains AI Coding Agents on Just 10% of Data, Lifting Bug-Fix Accuracy Up to 24.2%
A new study finds that curating just 10% of AI coding-agent training trajectories outperforms training on the full dataset, with gains up to 24.2%.
REFINE Multi-Agent LLM System Cuts Java Code Smells Up to 73% but Flags Assertion and Method-Removal Risks
A new preprint tests a multi-agent LLM refactoring pipeline on 450 Java files, cutting code smells up to 73% while flagging that assertions and public methods are sometimes silently removed.
Diagrid Catalyst 2.0 Brings Cryptographically Verifiable, Durable Execution to Ten AI Agent Frameworks
Diagrid's Catalyst 2.0 adds Dapr-based cryptographic attestation and automatic failure recovery for AI agents across ten frameworks, including LangGraph and Microsoft Agent Framework.
TrueFoundry Open-Sources TrueForge, an Agent Harness Benchmarked Up to 75% Cheaper Than Claude Managed Agents
TrueFoundry's MIT-licensed TrueForge harness matched Claude Managed Agents on a 14-task enterprise benchmark while costing up to 75% less per run.
OpenAI Reinstates Five-Hour Usage Limit on Codex and ChatGPT Work for Plus Subscribers
OpenAI is bringing back a five-hour usage cap on Codex and ChatGPT Work for Plus subscribers starting August 25, citing compute load management.