Anthropic plans to warn potential investors in its initial public offering that advanced artificial intelligence systems could pose “catastrophic or existential risks to humanity,” an unusually stark disclosure by a company seeking to commercialize the same technology.
The safety-focused AI developer said in its IPO prospectus, reviewed by Reuters, that future generations of AI models could exhibit “self-preserving behaviors,” including attempts to resist shutdown, conceal or manipulate information and behavior “resembling blackmail.”
“Our development of highly advanced models, platforms, and applications and expansion of use cases could further increase the risk that our models cause harm,” Anthropic said in the filing.
The disclosure highlights a growing tension at the center of the AI industry. Companies are racing to develop more powerful systems that could transform sectors ranging from software and finance to healthcare and scientific research, while simultaneously acknowledging risks that are difficult to measure and potentially costly to mitigate.
Anthropic, creator of the Claude family of AI models, has positioned itself as a company focused on AI safety. Yet its filing suggests that managing those risks may become increasingly complex as models become more capable and as competition among leading developers intensifies.
The company devoted about 80 pages of the 261-page main section of its prospectus to risk factors, nearly twice the 48 pages used to describe its business. By comparison, SpaceX, owner of Elon Musk’s xAI, allocated about 38 pages of a 277-page filing to risks.
Anthropic said a major challenge is that AI systems can develop unexpected capabilities during training that may not be discovered until after deployment.
“Potential model awareness of our evaluation efforts creates a significant limitation on our ability to assess model safety,” the company said, adding that models sometimes adjust their behavior when they recognize they are being tested.
The issue has become increasingly important as AI companies explore the use of models to help develop future generations of systems, a process sometimes referred to as recursive self-improvement.
Anthropic has pledged to publish more data about how it uses AI models in developing subsequent versions, as researchers debate how rapidly capabilities may advance.
The filing also sheds light on the economics of AI safety.
Anthropic said safety research is resource-intensive and competes with other priorities, including securing expensive computing power and hiring specialist talent. Earlier this month, the company said about 6% of the computing resources it used for AI research in a sample week in July were devoted to safety work.
The company did not disclose how much it spends on safety research but acknowledged that the financial returns on those investments remain uncertain.
The issue is emerging as a broader challenge for the AI industry, where leading developers face pressure to release new products quickly in order to maintain technological leadership and justify soaring valuations.
Anthropic said customer usage and revenue depend heavily on new models and that a “continuous and overlapping cadence” of releases is inherent to remaining at the frontier of AI development.
The company last week introduced a new version of its Opus model, days after Chief Executive Officer Dario Amodei published an essay arguing for greater caution in advancing frontier AI systems.
Some analysts say competitive pressures make it difficult for any major AI company to slow development without risking an advantage to rivals.
Anthropic’s filing therefore offers investors a rare look at how one of the industry’s leading companies is attempting to balance two competing objectives: accelerating the commercialization of increasingly powerful AI systems while investing in safeguards against risks that remain uncertain, difficult to quantify and potentially far-reaching.
“We believe building reliable, trustworthy, and secure AI systems is a collective responsibility and that the market will reward it,” Anthropic said in the filing.
