Meta and UIUC Researchers Get an 8-Billion-Parameter Model to Match Claude Opus 4.5 With a Smarter Agent Harness
EvoHarness-RL trains Qwen3-8B to hit 96.9% on ALFWorld, edging out Claude Opus 4.5's 96.4% baseline by rethinking agent memory and state.