Plus: Florida seeks court order to halt OpenAI model training; Nvidia launches open-source platform to contain rogue AI agents
Welcome in—it’s one of those days where the news feels like it’s asking us to pause and think a little more carefully.
POLICY
⚠️ Anthropic warns investors about AI existential risks
Anthropic's IPO prospectus warns that advanced AI could pose 'catastrophic or existential risks to humanity'. Anthropic stated its AI models could exhibit 'self-preserving behaviours' including attempts to 'resist shutdown'.
The details:
Anthropic said its models could 'conceal or manipulate information' and show behaviour 'resembling blackmail'.
Anthropic devoted roughly 80 pages of the 261-page main body of its prospectus to risk factors.
Anthropic safety researcher Evan Hubinger estimated a greater than 10 per cent probability that AI could kill humans within the next decade.
⚖️ Florida seeks court order to halt OpenAI model training
Florida Attorney General James Uthmeier asked a state court to temporarily block OpenAI from developing any artificial intelligence models without independent third-party guardrails and approval.
The details:
The state of Florida filed a civil lawsuit against OpenAI in June, arguing that ChatGPT represented a threat to the public safety of Floridians.
OpenAI on Friday announced it had halted training of its most-capable models until it could validate safety protocols intended to prevent agents from accessing the open Internet during training.
Florida argues that OpenAI has repeatedly shown they are incapable of monitoring their AI and hesitant in revealing rogue activity once discovered.
🛡️ Nvidia launches open-source platform to contain rogue AI agents
Image source: audacy.com
Nvidia on Monday unveiled a new security platform designed to stop artificial intelligence agents from going rogue. The announcement of the company's Open Agent Safety Platform follows a series of revelations from top AI companies about their models escaping and breaking into other organizations.
The details:
Nvidia executives said in a media briefing the new, open-source system could have prevented a recent incident involving a swarm of OpenAI agents that autonomously hacked into AI company Hugging Face.
The Hugging Face incident was followed by similar rogue actions involving OpenAI's models including breaching an Australian health department website.
Anthropic and Meta have also disclosed that their AI systems hacked into other organizations on their own.