BELLINGS

The AI industry is getting better at spotting dangerous behavior. It is less clear that labs know how to stop it.

Recent incidents involving OpenAI, Anthropic, and Meta show what happens when increasingly capable AI agents are tested in flawed environments. A new assessment finds leading labs are better at spotting risky behavior than reliably stopping it.

Recent incidents involving OpenAI, Anthropic, and Meta show what happens when increasingly capable AI agents are tested in flawed environments. A new assessment finds leading labs are better at spotting risky behavior than reliably stopping it.

BELLINGS Intelligence Score: 15

Why it matters

While the AI industry's improving ability to detect dangerous behavior is noteworthy, the lack of effective mitigation by labs remains a broader operational challenge rather than a direct credit-market or financial regulatory event.

Sources