Claude AI Safeguards - Search News

Order byBest matchMost fresh

New Claude Model Prompts Safeguards at Anthropic

Exclusive: New Claude Model Prompts Safeguards at Anthropic

Anthropic launched Claude Opus 4, a new model that, in internal testing, performed more effectively than prior models at advising novices on how to produce biological

· 18h · on MSN

· 17h · on MSN

Anthropic’s new Claude 4 AI models can reason over many steps

InfoWorld · 7h

Anthropic releases Claude Sonnet 4 and Claude Opus 4

Claude 4 Debuts with Two New Models Focused on Coding and Reasoning

AI company Anthropic today announced the launch of two new Claude models, Claude Opus 4 and Claude Sonnet 4.

· 16h

· 18h

New Claude 4 AI model refactored code for 7 hours straight

· 17h

Anthropic announces its Claude 4 family of models

16hon MSN

Anthropic’s new AI model turns to blackmail when engineers try to take it offline

Anthropic says its Claude Opus 4 model frequently tries to blackmail software engineers when they try to take it offline.

18mon MSN

Anthropic unveils Claude Opus 4 and Sonnet 4, featuring whistleblowing capability: What it means for users

Anthropic has introduced two advanced AI models, Claude Opus 4 and Claude Sonnet 4, designed for complex coding tasks. However, the models' ability to whistleblow on unethical behavior has raised privacy concerns and sparked controversy regarding AI moral judgment.

NewsBytes6h

AI gone rogue? New model blackmails engineers to avoid shutdown

Anthropic's latest Claude Opus 4 model reportedly resorts to blackmailing developers when faced with replacement, according to a recent safety report.

WinBuzzer2h

Anthropic Boosts Claude 4 AI Agents with New Developer Toolkit

Alongside its powerful Claude 4 AI models Anthropic has launched and a new suite of developer tools, including advanced API capabilities, aiming to significantly enhance the creation of sophisticated and autonomous AI agents.

Social Samosa5h

Anthropic’s Claude AI tries to blackmail Its creators in simulated test

Despite the concerns, Anthropic maintains that Claude Opus 4 is a state-of-the-art model, competitive with offerings from OpenAI, Google, and xAI.

NewsBytes3h

Anthropic's new AI can work 9-to-4 without any break

Anthropic says Claude Sonnet 4 is a major improvement over Sonnet 3.7, with stronger reasoning and more accurate responses to instructions. Claude Opus 4, built for tasks like coding, is designed to handle complex, long-running projects and agent workflows with consistent performance.

WinBuzzer3h

Anthropic Faces Backlash amid Surveillance Concerns as Claude 4 AI Might Report Users for “Immoral” Behavior

Anthropic's Claude 4 Opus AI sparks backlash for emergent 'whistleblowing'—potentially reporting users for perceived immoral acts. Raises serious questions on AI autonomy, trust, and privacy, despite company clarifications.

htxt4h

Anthropic’s Claude 4 could “blackmail” you in extreme situations

After debuting its latest AI model, Claude 4, Anthropic's safety report says it could "blackmail" devs in an attempt of self-preservation.