MUST READ | Trump hosts Anthropic CEO Dario Amodei amid AI safety dispute: What’s behind the meeting?
According to Reuters, which has reviewed the prospectus, Anthropic, led by Dario Amodei, has highlighted risks associated with its AI models, which it said could exhibit "self-preserving behaviors," including attempts to "resist shutdown," to "conceal or manipulate information" and behavior "resembling blackmail."
"Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm," Anthropic said in the filing.
DON'T MISS | US govt vs Anthropic: DC appeals court upholds Pentagon blacklist order. What it means?
Anthropic devoted 80 pages of its 261-page document on risk factors. "Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety," it said in the prospectus, adding that models sometimes develop unexpected capabilities during training that may not be discovered until they have been deployed and have resulted in significant safety incidents.
As per the report, AI researchers have warned that as models become more capable, they begin to understand and recognise when they are being watched. They then adjust their behaviour accordingly, making it harder to monitor them.
MUST READ | Anthropic introduces Claude Opus 5.5, its new lower-cost AI model for agentic coding, computer use, and more
Despite emphasising on AI safety, Anthropic said that the returns on its investments on safety are unclear. While it did not disclose how much it spent on AI research, Anthropic had earlier stated that 6 per cent of its computing power was dedicated to safety work.
"We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it," it said in the filing.
DON'T MISS | AI safety or market control? Anthropic, OpenAI, Google SpaceXAI accused of slowing AI progress
In recent weeks, Anthropic announced its commitment to sharing more data publicly regarding its use of AI models in developing future generations of the technology.