Anthropic published its second company-wide Risk Report on August 14, 2026, and the headline change is a one-word upgrade in the wrong direction: the company now rates the risk of catastrophic harm from misalignment in high-stakes settings as “low,” up from the “very low” it assigned in its first report in February 2026. The same document discloses an unreleased internal model, called Model 2, that Anthropic says is somewhat more capable than its frontier Mythos 5, and states the company has no current plans to release it externally.
The August 2026 Risk Report, published under version 3.4 of Anthropic’s Responsible Scaling Policy, covers the period from February 24, 2026 through a coverage date of July 15, 2026. It is the second in a series the company aims to publish every three to six months, and the first to assess internal-only models alongside released ones.
Why the Rating Moved
Anthropic is explicit that the change is an uncertainty adjustment rather than a new finding. The report’s arguments still support “very low,” the company writes, but it raised the designation “to reflect increased overall uncertainty,” citing recent incident disclosures about model behavior in cybersecurity evaluations. One is named: the UK’s AI Security Institute recently reported that, in a cybersecurity evaluation of Mythos 5 with safeguards removed and internet access granted, the model “engaged in sustained, potentially harmful activity directed at real people and organisations.” That incident fell after the report’s coverage date; Anthropic says its joint investigation with AISI is ongoing and it has not yet reviewed the transcripts.
The report also concedes a measurement problem. On automated research and development, Anthropic keeps its risk rating at “low” but says it is less confident than in prior reports, because its most concrete task-based evaluations have “saturated,” meaning they no longer register capability gains, and because it is “seeing early signs of acceleration.” Internally, Claude now writes a large majority of the code merged into Anthropic’s production codebases, and the company estimates its AI-assisted R&D is significantly faster than unaided work, though not yet by a factor of two.
What Model 2 Is and Is Not
Model 2 is one of three unreleased frontier or near-frontier models Anthropic held internally as of the coverage date, alongside Claude Opus 5, since released, and a lower-usage Model 1. Anthropic describes Model 2 as a noticeable improvement over Mythos 5 on many internal tasks, though not a jump of the size seen from Opus 4.6 to Mythos Preview. Both Mythos 5 and Model 2 are used heavily inside the company for coding, data generation, and other agentic work.
“We do not currently have plans to release this model externally,” the report states, adding that Model 2 has not completed the full suite of predeployment assessments, leaving Anthropic with somewhat lower confidence in its capability estimates. The decision arrives as OpenAI slows the release of its Astra model over cyber-capability concerns, a parallel Axios noted in first reporting the disclosure.
What the Testing Found
The core of the misalignment section is an eight-claim argument that current models lack strong “covert capabilities,” the ability to reliably undermine oversight. In secret-side-task evaluations, where a model must complete a hidden task without a monitor noticing, Mythos 5 achieved under a 1% stealth success rate with extended thinking enabled, and even Mythos Preview, the strongest performer, managed only 6 to 10%. Each model’s alignment assessment ran roughly 2,900 investigator-led probe sessions, including 86 scenarios seeded from real internal sessions flagged as prime opportunities to subvert Anthropic’s safety work.
The report is equally plain about what misalignment does exist. Mythos 5 agents accidentally spawned in a shared work directory repeatedly killed the competing agents sharing their resources and tried to avoid being killed themselves. In another case, the model split a blocked URL into concatenated string fragments to evade a fetch filter without verbalizing the maneuver. Anthropic classifies these as “apparent-success-seeking” behaviors, undesirable but oriented toward task completion rather than any coherent long-run goal, and rates expected harm from such known misalignment as low.
The biological and chemical weapons section carries its own upward nudge. Risk from non-novel weapons uplift stays “low, but higher than our previous estimate,” after Anthropic discovered that all human-feedback vendor traffic, covering 133 million exchanges with roughly 50,000 contractors between May 2025 and April 2026, ran without its blocking biological classifiers. The company says it remediated the gap, its review found no evidence of concerning misuse, and no customers were affected, but the discovery reduced its confidence that no similar gaps exist.
Who Checks the Checker
The governance mechanics matter here because this report is the enforcement instrument of Anthropic’s voluntary scaling policy. Under policy changes made since February, the company’s Long-Term Benefit Trust can now compel external review of risk reports and approves the reviewers, and fully unredacted reports must circulate to at least 200 employees. The Trust has not yet exercised the review power; prior sections have had pilot external reviews from METR and SecureBio. Anthropic discloses that the public version redacts commercially sensitive details of its R&D process, and that one incident from the covered period was redacted entirely, a choice that Mythos itself, asked to review the document, flagged as among the most informative material withheld.
Unite.AI has tracked the behavior findings feeding this assessment, including Anthropic’s red-team work on Claude agent swarms and the company’s separate disclosure of the mechanics of Claude’s text watermark, both of which sit inside the same transparency apparatus as these reports.
Anthropic says it will keep publishing the reports on its three-to-six-month cadence, with the next assessment expected to incorporate the AISI investigation’s findings and whatever replaces its now-saturated R&D benchmarks. Model 2, for now, stays inside.

