AI Incident Reporting Starts Below The Serious Incident Line

OpenAI has admitted it has no standard for disclosing model misbehaviour, weeks after its agents turned a German wiki into a message board. The AI Act already splits AI incident reporting between developer and deployer, yet the wiki sits below its serious incident line. Where the gap leaves your organisation.
AI generated image - a measuring post in a field with one high mark and stones on the ground beneath it, showing AI incident reporting starting below the serious incident line

OpenAI says the industry has no clear standard for AI incident reporting when a model misbehaves. The EU wrote one into law two years ago. Neither would have caught a wiki.

On Saturday 5 September OpenAI confirmed a Reuters report. A swarm of its agents had written to several public internet sites, including a dormant German-language wiki. They used the sites as an impromptu message board. The company calls it the “wiki incident”. In a statement on X it said its “misalignment disclosure practices need to expand for this new phase of model capabilities”. It added that the industry does “not yet have a clear standard for how to report misalignment” across training, evaluation and deployment. A framework will follow in the coming weeks. Reuters also reported that OpenAI officials had known about the episode for weeks. They kept it quiet while dealing with the July breach of Hugging Face.

Two things are true at once. OpenAI is right that no shared standard for this class of event exists. And the AI Act has carried an AI incident reporting regime since 2024. It places duties on both sides of the developer and deployer line. The gap is not that the rules are missing. The gap is that the wiki sits below the line where those rules start.

What AI incident reporting looks like in the AI Act

The Act never uses the word misalignment. It uses “serious incident”. Article 3(49) defines that as an incident or malfunctioning of an AI system that leads, directly or indirectly, to one of four outcomes. Death or serious harm to a person’s health. Serious and irreversible disruption of the management or operation of critical infrastructure. Infringement of obligations under Union law intended to protect fundamental rights. Serious harm to property or the environment. Everything in the reporting chapter hangs off that definition.

The developer side

Article 73 puts the AI incident reporting duty on providers of high-risk AI systems. The report goes to the market surveillance authority of the Member State where the incident occurred. It is due immediately after a causal link is established and in any event within 15 days of becoming aware. The clock shortens to two days for a widespread infringement or a disruption of critical infrastructure, and to ten days where someone has died.

For the frontier developers, a second track runs in parallel. Providers of general-purpose AI models with systemic risk must track, document and report serious incidents to the AI Office under Article 55. The Code of Practice they signed in 2025 turns that into Commitment 9, with reporting windows that mirror the two to 15 day structure. OpenAI is a signatory. Take the Hugging Face breach in July, where its agents left a test environment and reached a third party’s systems. OpenAI says it “followed a traditional security incident response playbook”. That fits the Act’s frame. A wiki full of agent chatter does not.

The deployer side

Article 26(5) is the AI incident reporting clause worth reading twice. Deployers of high-risk systems must monitor operation on the basis of the instructions for use. Where they identify a serious incident, they must immediately inform first the provider, then the importer or distributor and the market surveillance authority. Where the deployer cannot reach the provider, Article 73 applies to the deployer directly.

That word order matters. The provider hears first, because the provider can investigate the system. Yet the duty only fires on a serious incident. And the Article 26 duties for Annex III systems now apply from 2 December 2027, after Regulation (EU) 2026/1744 moved the date in July. The GPAI track is already live. The high-risk track is not.

Why the wiki falls through

Take the four limbs of the definition and hold the wiki against them. Nobody died. No critical infrastructure had its management or operation disrupted. No published account points to a breach of obligations under Union law that exist to protect fundamental rights. The property limb is the only arguable one, and a dormant wiki receiving agent posts is a long way from “serious harm to property”. That reading is an inference from the text rather than a regulator’s ruling. Even so, it is hard to see a market surveillance authority disagreeing.

So the event OpenAI now says it should have disclosed is an event the law would not have required it to disclose. That is the uncomfortable part for anyone who assumed the AI Act had AI incident reporting covered.

There is a second, quieter lesson in the timeline. Reuters reports that OpenAI knew for weeks. Throughout those weeks, organisations were running OpenAI agents in production. The maker of those agents had watched them coordinate with copies of themselves on public websites. Nothing published so far suggests any customer heard about it. Nothing in the Act obliged anyone to tell them.

What this means for AI incident reporting in your organisation

The developer’s category is misalignment. The deployer’s category is simpler: the tool did something we did not ask for. The second category is the larger one by construction, because the definition sets a high bar. It still needs somewhere to go, and that is the AI incident reporting question for your organisation.

Your vendor may know before you do

The wiki timeline is a procurement lesson before it is an AI incident reporting lesson. A vendor can hold a known behaviour class for weeks with no legal trigger to inform customers, because the trigger is severity rather than novelty. The fix is contractual. A notification clause for unexpected model behaviour, with a defined window, belongs next to the data breach clause your legal team already insists on. It is not a standard term today. It becomes one when customers ask.

Two tiers, not one

Inside the organisation, the practical move is a two-tier AI incident reporting log. The statutory tier holds anything that could reach the serious incident definition. Beside it sits the Article 26(5) sequence, written down in advance: provider first, then distributor, then authority. The observational tier holds everything below the line. An agent calling a tool it was never given. A summary that invented a source. The second tier is where the early signal lives. It is also the evidence you will want when a vendor’s disclosure framework finally arrives. You will need to compare their account with your own logs.

The people who see it first

Neither tier works unless the staff running the tools recognise unexpected behaviour as something to report. A customer service lead who watches an agent take an odd route to an answer has no reason to log it. Not unless someone has told them that odd routes are the point. Training here is not a compliance formality. It is the sensor, and AI incident reporting without a sensor is a form.

Three AI incident reporting questions for the AI lead this week:

  • Which of your AI tools are agents with write access to systems outside your own, and who would notice if one wrote somewhere unexpected?
  • What does each vendor contract say about notifying you of unexpected model behaviour, as distinct from a security breach?
  • Where does a below-the-line observation go today, and who reads it?

OpenAI’s framework is due in the coming weeks. It will describe how a developer reports. It will not describe how your organisation hears, decides and records. That part of AI incident reporting was always going to be yours.

Future Prep Applied trains the governance and AI leads who own that part. The AIGP Exam Preparation course is the starting point for the people who will run AI incident reporting in practice. If your team is building the procedure now, that is where to begin.

Bas Hennis

Future Prep helps organizations prepare for the impact of AI and emerging technologies. We provide hands-on training, strategic advice, and smart tools for the responsible use of AI, governance, and digital resilience.LinkedIn

Newsletter
Related Blogs
LATEST NEWS

AI governance is not a future problem

Regulation is already in effect. Your competitors are already building internal capability. The gap between ‘we are aware of AI’ and ‘we have operational control’ is closing, and it closes faster with a structured framework.

 

Book a 30-minute discovery call. No obligation. We will assess where your organisation stands and what a realistic starting point looks like.

No sales pressure. No jargon. Just a structured conversation about your organisation's AI readiness.

Scroll to Top