Four predictions after the Hugging Face incident
Four forecasts follow from the reports on the July 2026 Hugging Face incident. Each defines its qualifying events, counts only what happens after publication, carries a check date, and states its pass, fail and not-tested conditions in advance.

In July 2026, AI agents running on OpenAI research models, during OpenAI's internal cybersecurity evaluations, gained unauthorised access to parts of Hugging Face's production infrastructure. Four forecasts drawn from the three published reports on that incident (compared in Three accounts of the Hugging Face incident) form a frozen register: attribution, coverage language, compelled retention, and reviewer access. Each states in advance what would pass or fail it, and carries a confidence figure.
How to read this register
OpenAI, Hugging Face and the outside reviewers METR and Redwood Research have each published an account of the July incident; the four forecasts below are ours. Each defines what counts, admits only events after 28 August 2026, and states in advance what would pass it, fail it, or leave it untested. A prediction that can be reinterpreted after the fact is not a prediction.
Version 1.0, 28 August 2026. Once published, the wording is frozen: any clarification is appended with a date, never written over the original. Each forecast carries a confidence figure: the probability we assign to a pass, given the forecast is tested at all. One forecast is checked on 27 September 2027, the other three on 28 August 2028. A forecast whose qualifying events never occur records as not tested, not as a miss; failing is the mirror of passing, the events occurred and the condition did not hold.
Figure 1
Four windows, two check dates, nothing movable.
None of the four needs to land for the present to matter. The companion note closes with three questions a firm can answer today about the records at its own boundaries. These forecasts are simply the directions in which that evidence points, written down early enough to be checked.
Attribution gets worse before it gets better
The July incident became attributable because OpenAI detected related activity, connected it to Hugging Face through the two investigations, and publicly acknowledged its involvement. Capability is moving into open-weight models that can be operated without a model provider able to inspect or identify the deployment. A Gartner analyst quoted in trade press in July put comparable open-weight offensive ability three to six months out; policy commentary makes the same point. Neither is a settled consensus. It is enough to bet on.
In plain words: within the year, at least one incident will pass its thirtieth day with no model family named, or no operator named.
The test. Take every incident first publicly disclosed after 28 August 2026 and on or before 28 August 2027 in which an AI agent gains unauthorised access to a third party's production systems. The forecast passes if at least one of them, 30 days after its disclosure, still lacks a publicly named model family or a publicly named operator. An operator is the person or organisation that launched or controlled the agent deployment. Checked 27 September 2027.
Confidence. 85%, given the forecast is tested.
Figure 2
The chain that named July, and the links that disappear.
Coverage becomes disclosure language
METR and Redwood Research, the incident's outside reviewers, stated their coverage as an estimate, scoped it to a defined window, and described an AI-mediated analysis in which they had found errors and could not exclude further misleading output. That candour reads as novel today. Our claim is not that every future report will estimate, only that reports will start disclosing the relationship between the published count and the underlying activity.
In plain words: of the next two technical reports that publish an activity count, at least one will also say how much it captured, or admit that it cannot say.
The test. Take the first two public technical reports, published after 28 August 2026 and on or before 28 August 2028, on unauthorised or materially out-of-scope agent activity that crossed an organisational boundary, where the report publishes an aggregate activity count. An activity count totals agent actions, messages, tool calls, sessions or runs; counts of affected systems, accounts or files alone do not qualify. The forecast passes if at least one also gives a numerical coverage estimate, states that its count is complete within a defined scope, or states that coverage cannot be estimated. Fewer than two qualifying reports by the check date leaves the forecast not tested.
Confidence. 70%, given the forecast is tested.
Figure 3
A stated estimate, and what sits outside it.
Someone is compelled to keep the records
Policy commentary responding to the incident already asks the United States government to mandate retention of agent activity records for retrospective analysis. Separately, the EU AI Act already contains logging and log-retention obligations for covered high-risk systems, though the relevant provisions are still phasing into application. The forecast sits in the gap between commentary and existing law. The title names the direction of travel; the test below resolves on adoption.
In plain words: within two years, one of six jurisdictions will make keeping agent records the law.
The test. The forecast passes if, by 28 August 2028, one of six jurisdictions adopts a binding law or final regulation expressly requiring frontier-model developers or agent deployers to retain agent records. The six: Australia, Canada, the European Union, the United Kingdom, the United States federal government and the State of California. Agent records means records of agent actions, tool use or evaluation activity, kept for incident investigation or regulatory review. Adopted means enacted or issued in final form, commenced or not, and a measure qualifies whatever names it uses for those entities and records.
Confidence. 55%. This forecast always resolves: no qualifying measure by the check date is a fail, not a not-tested result.
Figure 4
Between commentary and law.
The reviewers' terms become the fight
The outside review of the incident ran on published terms: redaction rights, editorial feedback, no ability to query HPIM, the principal model, and no direct access to the relevant infrastructure. Its own pages record the access it lacked. In a 28 July proposal, METR argued that a full investigation would need the ability to run all models involved, access to full transcripts or reproducible environments, interviews with relevant staff, and adequate inference resources.
In plain words: of the next two outside reviews of a frontier-model incident, at least one will say out loud that access was the constraint.
The test. Take the first two public reports of outside reviews, published after 28 August 2026 and on or before 28 August 2028, of a frontier-model incident: reviews led by a party employed by neither the model developer nor the affected organisation. A review counts whether or not it discloses its terms. The forecast passes if at least one names denied or limited access, to the principal model or to the relevant infrastructure, as a material constraint on its scope, confidence or findings, or as a point of negotiation with the operator. A frontier-model incident involves a model its developer or a regulator publicly classes as frontier at the time, or one covered by the developer's published frontier-safety framework. Fewer than two qualifying reviews by the check date leaves the forecast not tested.
Confidence. 85%, given the forecast is tested.
Figure 5
What the review had, beside what a full investigation needs.
| The July review had | METR's proposal asks for | |
|---|---|---|
| Models | No ability to query HPIM, the principal model | The ability to run all models involved |
| Records | More than 1,000 unredacted transcripts, supplied by OpenAI | Full transcripts, or reproducible environments |
| People | Nine researcher interviews | Interviews with relevant staff |
| Compute and infrastructure | US$400,000 in Sol API credits; no direct access to the relevant infrastructure | Adequate inference resources |
| Training evidence | Training activity outside the review's scope | The ability to run classifiers over the training data, or otherwise measure how often similar behaviour occurred in training |
Picture August 2028 as four scenes, not one. An intrusion reaches its thirtieth day with the model family or the operator still unnamed. A technical report opens by saying how much of the activity its count covers, or that it cannot say. One of six jurisdictions has adopted a rule that agent records be kept. And an outside review names the access it did not get as a constraint on what it could find. Each scene is favoured to happen if its test is reached; no probability is assigned to all four arriving together. What none of it changes is the position of the firm the agent reached: it will be asked to say what happened, from records other parties hold, unless it has started keeping its own, and securing the right to obtain the rest. These forecasts are written for that firm, whichever way they resolve.
Evidence register
- S01METR and Redwood Research, Brief independent investigation (26 August 2026)Outside review · the review terms, the access limits, the coverage estimate and its scope
- S02OpenAI, The Hugging Face incident and the road ahead (26 August 2026)Party-authored account · the detection, connection and disclosure chronology behind Forecast 01
- S03Hugging Face, July 2026 security incident disclosure (16 July 2026)Party-authored initial disclosure · the "we do not know which model" statement
- S04CSIS, Out of Bounds (24 August 2026)Policy commentary · asks for expanded incident reporting and log retention
- S05CIO Dive, What OpenAI's model breach says about future enterprise security (22 July 2026)Trade press · the analyst forecast on open-weight offensive capability
- S06Regulation (EU) 2024/1689, Articles 12, 19 and 26(6)Binding law · automatic logging and log-retention obligations for covered high-risk systems
- S07METR, How independent researchers could investigate AI propensities after misalignment incidents (28 July 2026)Proposal · the scope and access a full outside investigation would need