Anthropic’s Model 2 Is Already at Work
The unreleased model is only about six weeks ahead on Anthropic’s aggregate capability trend. On the company’s internal R&D benchmark, the jump is considerably larger.
Anthropic has a frontier model that has no product page, no public API and no disclosed commercial name.
The company calls it Model 2.
As of the July 15 cutoff for Anthropic’s August 2026 Risk Report, Model 2 was an internal-only, Mythos-class system. Anthropic described it as somewhat more capable than Mythos 5, noticeably better on many tasks relevant to internal use, and slightly more capable overall—although stronger in some areas and weaker in others. It was being used internally at approximately the same level as Mythos 5. Anthropic said it had no current plans to release the model externally.
That final sentence has already attracted the obvious interpretation: Anthropic built a model too dangerous to release.
The report does not say that. It does not connect the absence of release plans to a failed safety evaluation, nor does it describe Model 2 as a dramatic capability discontinuity. Anthropic says the opposite on both points: the improvement is smaller than the jump from Opus 4.6 to Mythos Preview, and its internal approval process found no new or more concerning category of misalignment than had already been observed in Mythos 5.
The more interesting interpretation is operational.
Model 2 appears to be a model built far enough beyond Anthropic’s external product line to matter inside the company, but not so far beyond it that the model represents a new species of machine intelligence. Its aggregate benchmark gain is modest. Its gain on Anthropic’s own research and engineering problems is much larger. And it is already working inside
the organization through coding systems, data-generation pipelines and persistent agents.

What does 1.5 AECI points mean?
Anthropic gives one aggregate quantitative estimate for the model:
Model 2 is approximately 1.5 points above Mythos 5 on AECI.
AECI is Anthropic’s private version of Epoch AI’s Epoch Capabilities Index, or ECI. ECI attempts to solve a recurring problem in AI measurement: individual benchmarks become obsolete. Once every frontier model scores near 100%, the benchmark can no longer distinguish among them.
ECI combines results from many benchmarks and places models on a shared capability scale. It uses a statistical technique derived from item response theory, originally developed for educational testing. The model estimates three things simultaneously:
how capable each AI model is;
how difficult each benchmark is;
how sharply performance rises as models become capable enough to solve that benchmark.
This lets the index connect an older, easier test to a newer, harder test through models that took both. Epoch’s public ECI currently draws on more than 50 benchmarks. Anthropic’s AECI instead uses internal benchmark results, so its numbers cannot be compared directly with Epoch’s public leaderboard. Both scales use arbitrary units rather than percentages.
Anthropic’s chart fixes Claude Sonnet 3.5 at 130 AECI points and calculates a historical frontier trend of 13.5 AECI points per year through Opus 4.6. Model 2’s estimated gain over Mythos 5 is 1.5 points.
The conversion is straightforward:
That is:
1.33 months
40.6 days
approximately six weeks of progress at Anthropic’s earlier frontier trend
That is the correct way to contextualize the number.
It does not mean Model 2 is 1.5% more intelligent. It does not mean it was trained 41 days after Mythos 5. It does not mean every task improves by six weeks. It means that, if Anthropic had continued moving up its historical AECI trend line at 13.5 points per year, a 1.5-point gain would normally have taken about 41 days.
Anthropic also attaches large error bars to the estimate and does not plot Model 2 on the chart. It argues that comparisons between adjacent models may be somewhat more reliable than the absolute confidence intervals suggest because some uncertainty is shared across nearby models. Even so, this is a rough estimate, not a precision measurement.

The graph argues against the most sensational reading. Model 2 is not six months or two years beyond Mythos 5 on Anthropic’s aggregate measure. It is roughly six weeks ahead.
The next graph is more consequential.


