Safety & Ethics
31 articles RSS
Mistral Releases Shieldstral, a 3B Open-Weight Model That Moderates Text and Images Against Plain-Language Policies at Inference Time
Mistral AI open-sourced Shieldstral, a 3B-parameter safety classifier that judges content against natural-language policies without retraining, released as the first project from NVIDIA's Open Secure AI Alliance.
Red Hat Launches Asago, an Open Source Project to Automate AI Governance From Policy to Production
Red Hat and partners including NVIDIA, Microsoft, and IBM Research launched asago, an open source project that automates translating AI governance policy into deployable safety controls.
OpenAI's GPT-Red Cuts GPT-5.6 Sol's Prompt-Injection Failures Sixfold Through AI-on-AI Red-Teaming
OpenAI says its internal attacker model GPT-Red cut GPT-5.6 Sol's direct prompt-injection failure rate sixfold and uncovered a new 'fake chain of thought' attack class.
Anthropic's 'Cadences' Economic Index Links Its Survey to Real Claude Usage, Finding Heavier Automators More Optimistic About Their Jobs
Anthropic's June 2026 Economic Index links a ~9,700-person survey to actual Claude usage, finding people who delegate more work to AI are more optimistic about their jobs.
OpenAI's 'Deployment Simulation' Replays 1.3 Million Past Conversations Through Candidate Models to Predict Misbehavior Before Release
OpenAI's pre-release method regenerates real past conversations with an unreleased model to estimate undesired-behavior rates, with a median multiplicative error of 1.5x.
CDT Study Catalogs 37 'Dark Patterns' Across AI Chatbots, From ChatGPT and Claude to Replika and Character.AI
A Center for Democracy & Technology study built a taxonomy of 37 manipulative design patterns it found across major AI chatbots and companion apps.
Anthropic's Natural Language Autoencoders Turn Claude's Internal Activations Into Readable Text, Revealing Hidden Reasoning Patterns
A new Anthropic interpretability technique converts Claude's internal activations directly into plain-English descriptions, exposing evaluation awareness and reasoning the model never vocalizes.
Anthropic and the Gates Foundation Form a $200 Million Partnership to Deploy Claude in Global Health, Education, and Agriculture
The four-year commitment — described as the largest deal of its kind between an AI company and a global philanthropy — targets health services for 4.6 billion people in low-income countries.
OpenAI Rolls Out GPT-5.5-Cyber to Vetted Defenders, a Month After Mocking Anthropic's Mythos as 'Fear-Based Marketing'
OpenAI launched GPT-5.5-Cyber on May 7, 2026 to vetted security teams via its Trusted Access for Cyber program, after CEO Sam Altman publicly criticized Anthropic's restricted Mythos rollout.
OpenAI Replaces Its 2018 Charter Tone With Five Looser Principles, Hours Before the Musk Trial Opens
Sam Altman published a five-principle framework for OpenAI on April 26, dropping the 2018 charter's pledge to step aside for a safer competitor and reframing the company as AI infrastructure for humanity, just before jury selection in Elon Musk's $134 billion lawsuit.
Berkeley Researchers Hit Perfect Scores on Eight Top AI Agent Benchmarks Without Solving a Single Task
A UC Berkeley team showed that SWE-bench, GAIA, WebArena and five other widely cited agent benchmarks can be exploited to near-perfect scores, calling into question how the industry measures AI capability.
From Lab to Deployment: Mechanistic Interpretability Moves From Research Curiosity to AI Safety Tool
Anthropic, Google DeepMind, and OpenAI are integrating mechanistic interpretability into pre-deployment safety checks, marking a shift from academic technique to frontline defense.