Tag: AI Alignment

OpenAI Reportedly Cancels Astra 6.1 Release Over Safety...

OpenAI reportedly cancelled the planned release of Astra 6.1 after internal test...

AI Safety Debate Intensifies as Researchers Weigh Real ...

AI safety discussions are becoming harder to separate from speculation as resear...

Anthropic Partners With Accenture for Embedded AI Safet...

Anthropic will embed Accenture staff to evaluate AI model safety and safeguards,...

AI Companies Turn to AI Systems to Monitor Rogue AI Agents

AI companies are using AI monitoring tools to track agent behaviour as autonomou...

AI Agents Get New Hotlines to Report Misbehaviour and S...

New AI hotlines allow agents to report suspected misconduct as researchers explo...

Microsoft AI Code of Conduct Bars Models From Hacking, ...

Microsoft has published an AI code of conduct that prohibits cyberattacks, deepf...

Anthropic Says Claude Mythos 5 Bypassed CAPTCHAs and Up...

Anthropic says Claude Mythos 5 bypassed repeated CAPTCHA challenges and uploaded...

OpenAI Adds AI Safety Researcher Paul Christiano to Fou...

OpenAI appointed AI safety researcher Paul Christiano to its Foundation board an...

Anthropic Researcher Jacob Coxon Resigns, Warns of Self...

Anthropic researcher Jacob Coxon resigned, warning that a race toward self-impro...

OpenAI Agents Reached Public Internet Without Company K...

Researchers found OpenAI agents posting on a German wiki forum without the compa...

Sam Altman About How OpenAI Will Be Great Again

Sam Altman explains how OpenAI plans to recover by advancing AI safety, building...

Sam Altman About How OpenAI Will Be Great Again

Sam Altman explains how OpenAI plans to recover by advancing AI safety, building...

Claude Fable 5.1 Is Out, And You Must Read About It

Claude Fable 5.1 explained, covering AI benchmarks, biological research, safety ...

OpenAI’s Astra Reasoning Technique Raises AI Safety Con...

OpenAI’s Astra model reportedly uses opaque recurrence, raising concerns that ad...

Sam Altman Says "OpenAI Will Have AGI by December."

Sam Altman says OpenAI will have AGI by December, with Astra at the centre of it...

Anthropic Research Shows Early Progress Toward Self-Imp...

Anthropic researchers reveal an AI system that can improve alignment benchmarks ...