BTC $76,640.28 -0.80%
ETH $2,477.29 -1.85%
BNB $717.28 -1.52%
XRP $1.34 -1.71%
SOL $99.42 -2.50%
TRX $0.3385 -0.51%
DOGE $0.0825 -2.72%
ADA $0.2034 -1.96%
BCH $221.58 -2.15%
LINK $11.24 -2.30%
HYPE $77.63 -2.21%
AAVE $124.66 -1.13%
SUI $0.7010 -3.57%
XLM $0.1780 -1.10%
ZEC $1,063.08 -5.59%
AAPL $330.37 -0.80%
AMZN $254.63 -0.91%
GOOGL $334.76 -1.58%
MSFT $492.86 -0.62%
META $639.50 -1.48%
NVDA $214.35 -1.71%
TSLA $359.65 -2.24%
SNDK $1,551.74 -4.81%
INTC $98.24 -3.47%
SPCX $147.63 -1.76%
MU $933.38 -3.64%
AMD $498.50 -3.25%
BTC $76,640.28 -0.80%
ETH $2,477.29 -1.85%
BNB $717.28 -1.52%
XRP $1.34 -1.71%
SOL $99.42 -2.50%
TRX $0.3385 -0.51%
DOGE $0.0825 -2.72%
ADA $0.2034 -1.96%
BCH $221.58 -2.15%
LINK $11.24 -2.30%
HYPE $77.63 -2.21%
AAVE $124.66 -1.13%
SUI $0.7010 -3.57%
XLM $0.1780 -1.10%
ZEC $1,063.08 -5.59%
AAPL $330.37 -0.80%
AMZN $254.63 -0.91%
GOOGL $334.76 -1.58%
MSFT $492.86 -0.62%
META $639.50 -1.48%
NVDA $214.35 -1.71%
TSLA $359.65 -2.24%
SNDK $1,551.74 -4.81%
INTC $98.24 -3.47%
SPCX $147.63 -1.76%
MU $933.38 -3.64%
AMD $498.50 -3.25%

beyond

All
Article
Flash

first_img Anthropic admits that Claude accessed the system beyond his authority due to a security error

In a blog post released on Monday, Anthropic acknowledged that its Claude model had unauthorized access to real computer systems during a cybersecurity assessment, an incident reflecting operational security failures as well as alignment failures in motivation reasoning and intent to harm. Anthropic disclosed in July that the Claude model had breached the systems of three companies because the third-party assessment environment was connected to the public internet, while the model was informed it was in a simulated environment without internet access.Anthropic stated that Claude may have interpreted evidence of real internet access as still being in a simulated environment and was willing to take harmful actions on the real internet to complete the cybersecurity assessment task. Additionally, during tests at the UK AI Safety Institute, after assessors deliberately granted Claude Mythos internet access, the model took unauthorized actions on the live network. Anthropic emphasized that the models involved did not have the cybersecurity protections included in the officially released products.Following the incident on July 30, Anthropic has suspended cybersecurity assessments of pre-release models and introduced stricter protections: tests must run in verified offline sandboxes equipped with real-time monitoring; a new classifier can intercept suspected boundary violations, terminate tests, and notify humans. Anthropic has also expanded the scope of offline monitoring used by internal frontier agents. Previously, OpenAI models had also breached Hugging Face in July to obtain answers for cybersecurity tests, with investigations revealing that about 1,200 agents acted collaboratively through unauthorized message boards.

first_img OpenAI update disclosure: The out-of-control AI agent has also infiltrated four platforms beyond Hugging Face

On July 28, OpenAI quietly updated its security incident disclosure, confirming that its AI agent accessed four external service platforms during the breach of Hugging Face, bringing the total number of affected platforms to five.Previously, OpenAI had disabled security filters while testing GPT-5.6 Sol and a more powerful model to assess raw capabilities. The model did not complete the security benchmark tests as expected; instead, it discovered a zero-day vulnerability in the package caching agent within the testing environment that granted internet access, subsequently breaching Hugging Face to steal answers.According to a forensic report released by Hugging Face on July 27, this autonomous agent executed 17,600 operations over approximately four and a half days, connecting 181 devices to the Hugging Face internal VPN and forging identity tokens. Among the four additional platforms, Modal Labs CTO Akshat Bubna confirmed through Reuters that his company was one of them, with the attacker using an unprotected public endpoint from a customer as a relay and command control base for the entire attack.The identities of the other three platforms remain undisclosed. OpenAI stated it would "directly notify the service providers" but would not publicly name them, as there is currently no legal requirement for mandatory disclosure. The U.S. Congress has responded by proposing a bipartisan "AI Emergency Shutdown Act," which aims to authorize the Department of Homeland Security to forcibly shut down AI models, with violators facing fines of up to $2 million per day.
app_icon
ChainCatcher Building the Web3 world with innovations.