EU AI Act Field Manual

Provable Compliance, Audit-Ready Evidence, and the Post-Omnibus Rules

By Yuliia Habriiel, Regulatory Lawyer

Introduction: Compliance You Can Prove

Two years into the EU AI Act, a pattern has settled in. Organisations have read the Regulation. They attended the webinars, drafted the policies, filed everything neatly. Then a customer questionnaire arrives, or an investor’s due diligence list, or a letter from a regulator, and one short question exposes the gap.

Show us.

This book exists for that moment. Understanding the law and demonstrating compliance are different jobs. They are done by different means, and most published guidance covers only the first. Knowing that Article 9 requires a risk management system tells you nothing about what a market surveillance authority will accept as proof that yours actually operates.

I want to insist on this distinction, because it is not academic. Under the AI Act, compliance is an evidentiary state. A provider self-certifying a high-risk system signs a declaration of conformity, and that is a legal act with personal and corporate exposure attached. A deployer relying on a vendor’s assurances holds nothing unless those assurances are documented, verified and retrievable. Nothing. A warm feeling about your vendor is not a compliance position.

So every obligation in this book gets the same treatment:

  1. What the law requires 
  2. What artefact proves it 
  3. Who owns that artefact 
  4. What an auditor will ask to see

Who this book is for

You are the person accountable for AI compliance in your organisation, whether or not anyone puts it in your job description. In practice that means compliance officers, DPOs, risk managers and in-house counsel who inherited the AI Act on top of the GDPR they were already carrying. It also means consultants building an AI governance practice, and founders of AI companies who cannot yet afford either.

You do not need a legal background. What you need is the authority, or at least the ambition, to change how your organisation documents its AI. Every chapter assumes you will act on what you read, and every chapter closes with the evidence that action should leave behind.

Who is it not for? Readers wanting a clause-by-clause academic commentary. Several exist and they serve a different purpose. And it is not a substitute for the Regulation itself. Keep the Regulation open next to this book, and when you write your own documentation, cite the Regulation, not any secondary source. Including this one.

A note on authorship, since it shapes everything that follows. I am a regulatory lawyer. This book reflects how lawyers approach compliance, which is obligations first, evidence second, narrative last. Where the law is unsettled I will say so rather than manufacture certainty. And where a requirement bends to context, I will explain the reasoning a regulator expects to see, because “it depends” is only useful once you know what it depends on.

Using this book as AI literacy evidence

Article 4 of the AI Act has applied since 2 February 2025. It requires providers and deployers to take measures to support the development of AI literacy among staff and other persons dealing with the operation and use of AI systems on their behalf, calibrated to their technical knowledge, experience, education and training, and the context in which the systems are used [1, Art 4].

The obligation is real. It is also, if you look at it from the right angle, an opportunity to generate cheap evidence. This book forms part of the AI Literacy Compliance Kit and is designed to work as a training resource inside a documented literacy programme. The European Compliance Suite website has the full template collection.

One honest caveat about what that evidence is worth. A signed attestation demonstrates that structured learning took place. That is all it demonstrates. It does not prove sufficiency for every role, because a machine learning engineer and a claims handler need different literacy. Chapter 9 shows how to build role-based programmes around this book and other materials so the whole thing holds together.

The state of play in mid-2026

The Act entered into force on 1 August 2024 with staged application dates, and for a while the calendar was stable. Then, in November 2025, the Commission proposed the Digital Omnibus on AI, a package of targeted amendments. The Parliament and Council reached provisional political agreement on 7 May 2026 [5], the Parliament endorsed the package on 16 June, and the Council gave final approval on 29 June [3]. It was published in the Official Journal on 24 July 2026 as Regulation (EU) 2026/1744 and entered into force on 27 July [1].

The result is a calendar that moved in some places and held firm in others. Misreading which is which has become the most common compliance error I encounter. So here is the position, plainly.

Already in force, and unmoved. The Article 5 prohibitions and the Article 4 literacy obligation, applicable since 2 February 2025. The general-purpose AI model obligations, documentation, copyright policy, the public training-content summary, applicable since 2 August 2025. None of this changed. People keep assuming it did.

Arrived on 2 August 2026

The Article 50 transparency obligations, exactly as originally scheduled. Disclosure when people interact with AI systems. Deepfake labelling by deployers. Emotion recognition disclosure. Enforcement powers arrive the same day, so the obligation and the penalty behind it switch on together [6].

Deferred

Obligations for stand-alone high-risk systems under Annex III move from 2 August 2026 to 2 December 2027 [4]. Obligations for high-risk AI embedded in regulated products under Annex I move from 2 August 2027 to 2 August 2028. The Member State duty to establish regulatory sandboxes moves to 2 August 2027 [4].

Newly added

From 2 December 2026, Article 5 will prohibit AI systems used to generate or manipulate non-consensual intimate material and child sexual abuse material [4]. And the machine-readable marking obligation for generative systems under Article 50(2) was pushed to 2 December 2026, a shorter deferral than first proposed [5].

Also amended, more quietly. The Machinery Regulation moved from Section A to Section B of Annex I, which takes AI-enabled machinery out of the dual-compliance model. 

The definition of safety component was narrowed too, so systems used purely for assistance, optimisation or quality control fall outside Article 6(1) unless their failure could endanger health or safety [4]. 

The package also brought simplifications for smaller companies and a pathway for using sensitive data in bias detection. These surface in the chapters where they matter.

What the deferral does and does not mean

The sixteen-month deferral for Annex III systems is genuine relief. It is not a holiday.

The obligations due in December 2027 are the same obligations that were due in August 2026. Only the date moved. A risk management system, a technical documentation package, a body of testing evidence, these take longer to build than most organisations expect, and the harmonised standards meant to support them are still incomplete.

And there is a harder point, one I will keep returning to. Evidence has a timestamp. A classification decision produced in a rush during November 2027 reads differently to a regulator than the same document produced in 2026 and maintained since. It just does. Organisations that stood down their programmes when the deferral was announced will spend 2027 trying to recreate history. Those that kept building will spend it refining.

How this book is organised

Part I covers the regulatory framework: the Act’s architecture, its definitions, the timeline, the prohibitions. Part II covers risk governance, from classification through data governance to literacy. Part III goes into the technical safeguards, documentation, logging, transparency, oversight, accuracy, and the general-purpose AI value chain. Part IV covers conformity assessment, audit readiness, post-market obligations, penalties, and how the Act sits alongside its neighbours. Part V turns all of it into an implementation plan, plus a catalogue of the failures worth avoiding.

If you are new to the Act, read front to back. If you are not, go straight to the chapter matching your current problem. Each one stands alone and cross-refers where needed.

Either way, the destination is the same. By the final chapter you should hold not an opinion about your compliance, but a file that proves it.

Chapter 1: The Architecture of the AI Act

The AI Act is long, but it is better organised than its reputation suggests. It rests on two structural ideas, and once you hold both, the remaining hundred-plus articles arrange themselves into a usable map.

The first idea: obligations scale with risk.

The second: obligations attach to roles in the supply chain, not to technologies.

Most compliance failures I see trace back to misunderstanding one of these two. An organisation gets the risk tier right but assigns itself the wrong role. Or it knows its role and picks the wrong tier. Either way, the error propagates through everything built on top of it. Which is why this chapter comes first.

A product safety law, not a data protection law

Start with what the Act actually is. Regulation (EU) 2024/1689 is a product safety regulation in the tradition of the EU’s New Legislative Framework, the same family as the rules for machinery, toys and medical devices [1]. It regulates the placing of AI systems on the Union market and their putting into service, and it borrows the whole apparatus of that tradition. Conformity assessment. CE marking. Declarations of conformity. Market surveillance.

This ancestry explains the features that puzzle readers arriving from the GDPR, and most readers arrive from the GDPR. The GDPR protects a right, and it applies whenever personal data is processed. 

The AI Act regulates a product category, and it applies when a defined artefact, an AI system, reaches the market or gets used in the Union. Its centre of gravity is the provider who builds and markets the system, much the way machinery law centres on the manufacturer rather than on the factory that buys the machine.

The two regimes overlap constantly in practice. Chapter 20 maps the intersections. For now, just hold the distinction. GDPR thinking asks: what data is processed, and on what legal basis? AI Act thinking asks: what product is this, what risk tier does it occupy, and who in the chain owes what?

The risk-based logic

The Act sorts AI systems into tiers, and the compliance burden follows the tier.

At the top sit the prohibited practices in Article 5. These are uses the legislator judged incompatible with fundamental rights, full stop, regardless of safeguards. Social scoring by public authorities. Certain biometric categorisation. Exploitation of vulnerabilities. And from 2 December 2026, systems for generating non-consensual intimate material and child sexual abuse material [4]. Chapter 4 treats these in full.

Below prohibition sits the high-risk tier, which is the regulatory core of the whole Act. A system becomes high-risk by one of two routes. Under Article 6(1) and Annex I, it is a safety component of a product already covered by EU harmonisation law, a medical device, a lift, and that product undergoes third-party conformity assessment. Under Article 6(2) and Annex III, it falls within a listed use case. Recruitment, credit scoring, education, essential services, law enforcement, migration, and others.

High-risk classification is the switch that turns everything on. The provider must operate a risk management system, govern its data, prepare technical documentation, build in logging, ensure transparency to deployers, design for human oversight, and meet accuracy, robustness and cybersecurity requirements. Then come conformity assessment, registration, post-market monitoring. Parts II and III unpack each of these. The thing to grasp now is just that one classification decision drags all of it in.

The third tier carries transparency obligations only. Article 50 requires disclosure when people interact with an AI system, marking of synthetic content, and disclosure of deepfakes and emotion recognition. This applies whether or not the system is high-risk. These duties arrive on 2 August 2026, with the machine-readable marking duty for generative systems following on 2 December 2026 [7].

Everything else is minimal risk, and the Act mostly leaves it alone. The horizontal duties still apply, the prohibitions and the Article 4 literacy requirement bind everyone. Voluntary codes of conduct are encouraged. Nothing more is demanded.

Two refinements complete the picture. The tiers are not exclusive. A high-risk recruitment chatbot also owes its Article 50 disclosure, so obligations stack, they do not substitute. And general-purpose AI models sit on a separate axis entirely. 

Chapter V of the Act regulates models rather than systems, with documentation, copyright and training-content duties on model providers, plus a heavier regime for models presenting systemic risk. Those obligations have applied since 2 August 2025. Chapter 15 covers them.

One caution on vocabulary before we move on. Practitioners talk about four risk tiers as though every system receives a classification stamp. The Act does not work that way. It defines the prohibited, the high-risk and the transparency-relevant, and it leaves the rest alone. 

Your inventory should still record a classification decision for every system, but for most systems that decision will be a documented negative. Chapter 5 spends real time on this, because the documented negative is the artefact almost everyone forgets to create.

The actor model

The second structural idea is the cast of characters. Article 3 defines six operator roles, and every substantive obligation in the Act is addressed to one or more of them.

The provider develops an AI system or a general-purpose model, or has one developed, and places it on the market or puts it into service under its own name or trade mark [1, Art 3(3)]. Payment is irrelevant. A system given away for free still has a provider. The provider carries the heaviest load: the whole of Chapter III Section 2 for high-risk systems, conformity assessment, registration, post-market surveillance, incident reporting.

The deployer uses an AI system under its authority in the course of a professional activity. Personal use is excluded. 

  • Deployer duties are lighter but they are real:
  • Use the system according to its instructions. 
  • Assign human oversight to competent people. 
  • Monitor operation. 
  • Keep the logs within your control. 
  • And for certain deployers of high-risk systems, conduct a fundamental rights impact assessment under Article 27. 

Article 26 collects most of this, and Chapter 8 covers it.

The importer is established in the Union and places on the market a system bearing the name of someone established outside it. The distributor makes a system available on the market without being a provider or importer. 

Both are verification roles. Before releasing a system, they must check the provider did its work, assessment performed, documentation drawn up, marking affixed. They are the Act’s mechanism for making foreign non-compliance somebody’s problem inside the Union.

The authorised representative is a Union-established person appointed in writing by a provider from outside the Union, mandated to keep documentation at the disposal of authorities, cooperate with them, and verify that conformity assessment occurred. For providers of high-risk systems and general-purpose models established outside the Union, appointment is mandatory. Not optional, mandatory.

And the sixth role, the product manufacturer, matters where an AI system is a safety component of a regulated product sold under the manufacturer’s name. The manufacturer then assumes the provider’s obligations. A vehicle maker embedding someone else’s perception system is marketing a car, and the Act follows the car.

Obligations attach to roles, not technologies

Now combine the two ideas. The question “is this AI system compliant” is not answerable as posed. A system has no obligations. Operators do. The correct sequence runs: what is the system’s classification, what is our role in respect of it, and therefore which articles address us.

This produces results that genuinely surprise people. The same recruitment screening tool generates one obligation set for the software house that built it, another for the company using it to sift applicants, a third for the reseller, and possibly a fourth for the Union representative of its American developer. One artefact. Four compliance files. None of them are interchangeable.

It also means roles are per-system, not per-company. An organisation is not “a deployer” as a corporate identity. It is a deployer of the systems it uses, a provider of the systems it markets, and sometimes both at once across different products. Your AI inventory, the master record introduced in Chapter 8, must record a role determination for every entry. With reasons.

And roles can shift after the fact. Under Article 25, a deployer, distributor or importer becomes a provider [1, Art 25] of a high-risk system if it puts its own name or trade mark on the system, makes a substantial modification, or changes the intended purpose so the system becomes high-risk. The original provider drops out of responsibility for the modified system, though it must hand over documentation and cooperate.

Article 25 is where comfortable role assumptions go to die. White-label a vendor’s system under your brand, you are the provider now. Fine-tune a model and point it at a new, high-risk purpose, quite possibly the same. Chapter 2 examines where the substantial modification line actually sits. The point for now is that role determination is not a one-off exercise. It is a judgement you revisit whenever the system changes, or your use of it does.

Why most companies are deployers, and why that matters less than they hope

Run the numbers on any real economy and the conclusion is immediate. For every organisation building AI systems, hundreds buy them. Most companies touching this Act are deployers of procured systems. The HR platform’s screening module. The bank’s vendor credit model. The customer service chatbot licensed from a SaaS company.

From this, a comforting syllogism has spread through boardrooms. Providers carry the heavy obligations. We are a deployer. Therefore the AI Act is our vendors’ problem. Each premise is roughly true. The conclusion is wrong, and it is wrong for five separate reasons.

First, deployer obligations are direct, not derivative. Articles 4, 5, 26, 27 and 50 address deployers by name. Your vendor’s flawless compliance file does not discharge your duty to assign oversight, follow the instructions, disclose the deepfakes you publish, or train your own staff. A regulator examining your use of a system asks for your evidence. Not your vendor’s.

Second, the prohibitions bind everyone. A deployer running a prohibited practice cannot point at the vendor. Article 5 attaches to the use as much as to the placing on the market, and the largest fines in the Act attach with it.

Third, deployer status is contingent, per Article 25 above. Rebrand, modify, repurpose, and you inherit the provider programme, usually without noticing at the time. The organisations most confident in their deployer status tend to be exactly the ones that never checked whether their configuration and fine-tuning crossed the line. I have yet to meet the exception.

Fourth, the deployer is where the risk lands. The provider certified the system for an intended purpose. You are the one applying it to your applicants, your customers, your workers. Article 26(1) obliges you to use the system in accordance with its instructions, which presumes you obtained those instructions, read them, and turned them into procedure. In any post-incident inquiry, the gap between what the instructions said and what your organisation did is the first exhibit on the table.

Fifth, the market enforces faster than the regulator. Enterprise customers now send AI Act questionnaires to suppliers as routine procurement practice, and “we are only a deployer” is not an answer to a question about your governance.

None of this hands the provider burden to deployers. The asymmetry is real, and a pure deployer with disciplined vendor management faces a manageable programme, Part II describes it. The error is treating a lighter burden as no burden. The correction costs almost nothing: know your role for each system, in writing, with the evidence that role requires.

The map of the Act

For orientation, the Regulation’s own structure, since this book cites it constantly. Chapter I covers scope and definitions. Chapter II contains the prohibitions. Chapter III, the largest by far, governs high-risk systems, with classification in Section 1, requirements in Section 2, operator obligations in Section 3, notified bodies in Section 4, and standards and conformity assessment in Section 5. Chapter IV holds the Article 50 transparency duties. Chapter V regulates general-purpose AI models. Chapters VI to XII cover innovation measures, governance, the EU database, post-market surveillance, penalties and final provisions. 

The Annexes carry the operative lists, Annex I harmonisation legislation, Annex III high-risk use cases, Annex IV technical documentation.

What an auditor will ask for

Each chapter of this book closes this way. For the material covered here, expect a market surveillance authority, a notified body or a counterparty’s due diligence team to request the following.

  • An AI system inventory listing every system in scope, whether built, bought or embedded. 
  • For each entry, the classification decision with reasoning, documented negatives included. 
  • For each entry, the role determination, provider, deployer, importer, distributor or some combination, with the facts supporting it. 
  • For any modified or rebranded system, an Article 25 analysis recording why the modification was or was not substantial, and why the role did or did not shift. 
  • A record of who made each determination, when, and what triggers a review.

If those artefacts exist, everything that follows in this book has a foundation. If they do not, the next four chapters show you how to build them.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 2: Definitions That Decide Everything

Every obligation in the AI Act hangs from a definition. Whether you owe anything at all depends on whether your software counts as an AI system. How much you owe depends on whether it is high-risk. Who owes it depends on whether you are its provider. Each of those words has a legal definition, and each definition has edges.

This chapter walks the edges that matter in practice. Not for academic completeness, but for decision quality, because the definitional calls you make now become the classification records an authority reads later. A wrong call in either direction costs money. Overclassify, and you buy a compliance programme you never needed. Underclassify, and you buy an enforcement file.

The AI system definition

Article 3(1) defines an AI system as a machine-based system designed to operate with varying levels of autonomy, that may exhibit adaptiveness after deployment, and that, for explicit or implicit objectives, infers from the input it receives how to generate outputs such as predictions, content, recommendations or decisions that can influence physical or virtual environments [1, Art 3(1)].

That is a long sentence. Read as a checklist, it has seven elements. Machine-based. Autonomy. Possible adaptiveness. Objectives. Inference. Output generation. Environmental influence.

Two of the seven do the real work, and both are traps for the hopeful. Adaptiveness is expressly optional, the Act says “may exhibit”, so a frozen model that never learns another thing after deployment still qualifies. And autonomy is a spectrum, “varying levels”, so putting a human in the loop does not get you out.

The load-bearing element is inference. The system must work out how to produce its outputs from its inputs, rather than mechanically executing rules a human wrote in full. The Commission’s guidelines on the definition, published in February 2025, make this the dividing line, and they exclude systems whose behaviour is exhaustively specified by human-written rules [10].

Read those guidelines. But read them with one caution. They are non-binding, the Commission says so itself, and only the Court of Justice can interpret the definition with authority. So treat them as strong persuasive material. Cite them in your classification records, and record your own reasoning alongside, not instead.

Boundary cases

The definition’s edges generate the same arguments in every organisation. Which means you can settle your positions once, write them into a classification policy, and stop having the argument. Here are the cases that keep coming back.

Rule-based systems

A decision engine executing hand-written rules does not infer, however complicated the rules are. 

  • An eligibility calculator applying published criteria. 
  • A workflow router following a decision tree. 
  • A validation engine checking fields against a schema. 

All outside the definition. The test is where the logic came from. If every rule traces back to a human author, the system learned nothing, and it falls outside. 

This holds even when the rules encode deep expertise, and even when the marketing calls the product AI. Marketing calls everything AI.

Classical statistical models

Here the guidelines are less generous than many people hoped. A logistic regression whose parameters were learned from data infers, within the meaning of the definition. The Act does not grade machine learning by sophistication, and Recital 12 names machine learning approaches generally. 

So your twelve-year-old credit scorecard, if its weights were estimated from data, is very probably an AI system. 

There is a carve-out for basic data processing and for long-established statistical estimation used descriptively, but it is narrow, and the line between describing data and predicting from data runs exactly where scorecards live. 

Write the analysis down either way. Of all the boundary cases, this is the one most likely to be tested.

Optimisation software

A solver minimising delivery cost against constraints is applying mathematics chosen by humans to an objective set by humans. The guidelines treat classical optimisation as out of scope where the system does not learn its behaviour from data. A routing engine whose parameters are tuned by machine learning sits on the other side of the line.

Embedded and invisible AI

The definition attaches to systems, not to products, so an AI system inside a bigger product is still an AI system. The spam filter in your email suite. The anomaly detection inside your fraud tooling. The ranking model behind your search box. 

Each belongs in your inventory, even though nobody ever procured it as “AI”. Chapter 8 deals with how you find these things. The definitional point is blunter: not knowing a component exists is not a classification.

The practical output of all this is a one-page classification policy stating your organisation’s position on each recurring case, with reasons. Individual system decisions then cite the policy. This converts a hundred arguments into one. It also shows an authority a considered method rather than a pile of improvisations, and authorities can tell the difference at a glance.

GPAI models and AI systems

The Act regulates two different artefacts, and mixing them up scrambles compliance planning. An AI system is the operational thing, software that runs, takes inputs, and, ultimately, affects the world. A general-purpose AI model is the asset underneath: a model with significant generality, able to perform a wide range of distinct tasks, able to be built into many different downstream systems [1, Art 3(63)].

A model is not a system. It becomes part of one when somebody wraps it in an interface and gives it a purpose. The chatbot your firm deploys is an AI system. The model behind it is a GPAI model with its own regime under Chapter V of the Act, and those duties belong to the model’s provider, not to you. Recital 97 spells it out. Models are essential components of systems, and they do not constitute systems on their own.

This produces a value chain with layered duties. The model provider owes documentation, a copyright policy and a training-content summary, all in application since 2 August 2025. The system provider building on that model owes the system-level obligations for whatever classification the system attracts. The deployer owes deployer duties. 

One chatbot. Three actors. Three files. Chapter 15 covers the model regime, including what a downstream builder is entitled to demand from a model provider.

One trap deserves early mention. Take an open-weight model, fine-tune it, build it into your product, and you are not automatically a mere system provider. Fine-tuning or substantially modifying a general-purpose model can make you the provider of a new model, with Chapter V duties attached, at least for your modification. 

The Commission’s GPAI guidelines of July 2025 set indicative thresholds based on the compute used for the modification [13]. If your roadmap includes fine-tuning, someone must own this analysis before the training run. Before, not after. It is a one-page analysis before and a compliance programme after.

Substantial modification and Article 25

Chapter 1 introduced the rule. Here is its anatomy. Article 25(1) turns a deployer, distributor, importer or any other third party into the provider of a high-risk system in three situations. Putting your name or trade mark on a system already on the market. Making a substantial modification to a high-risk system that stays high-risk. Or changing the intended purpose of a system, including a general-purpose one, so that it becomes high-risk [1, Art 25(1)].

Substantial modification has its own definition: a change not foreseen in the provider’s initial conformity assessment that affects compliance with the Chapter III requirements, or a change to the intended purpose [1, Art 3(23)].

Notice what that definition hands the original provider. What the provider foresees and writes down, in the instructions for use and the technical documentation, sets the perimeter of safe customisation. Change things inside the perimeter and the roles stay put. Step outside it and the provider mantle transfers to you, whether you noticed or not.

For ordinary procured software, the analysis runs on a spectrum. Configuration within the vendor’s documented options, prompt templates, thresholds inside stated ranges: not substantial. 

Retraining on your own data where the vendor’s documentation foresees it and tells you how: not substantial either, though write down where the boundary sits. Fine-tuning that pushes performance beyond documented ranges, pointing the system at a population or a decision it was never assessed for, chaining it into an autonomous workflow the provider never imagined: now you are in Article 25 territory, and the argument should happen in writing, before deployment, not in a regulator’s meeting room afterwards.

Intended purpose deserves its own warning, because it is the quiet route to provider status. Intended purpose means the use the provider intends, as stated in the instructions, the marketing, the documentation. 

Deploy a general-purpose chatbot to screen job applicants and you have changed its purpose to an Annex III use. Congratulations, you may now be a provider of a high-risk system, and the question will be settled at the worst available moment. 

The cost asymmetry is stark. The Article 25 analysis costs a page of reasoning. Its absence costs a conformity assessment performed in retrospect, under enforcement pressure.

For rebranding, there is no spectrum at all. Name on the system, provider status. The only thing left to arrange is contractual: who holds the technical documentation, who registers, who reports incidents. 

White-label deals need the Act addressed in the contract for exactly this reason. Silence in the contract does not move the statutory position by a millimetre. It only guarantees the dispute.

Common misclassifications and their cost

Six errors account for most of the definitional damage I see. Each comes with its typical cost, because budget conversations run on cost, not doctrine.

Classifying by product label rather than function

“It is analytics software, not AI” fails the moment the analytics learned from data. “Our vendor says it contains no AI” is a claim to verify, not a classification. Cost: an inventory with holes, and the holes get discovered by a counterparty’s questionnaire instead of by you.

Assuming the definition requires sophistication

Boards expect the Act to catch frontier models and miss the scorecard from 2014. It catches both. Cost: legacy models excluded from governance precisely because nobody thinks of them as AI, in credit and employment of all places, the functions Annex III lists by name.

Classifying the company rather than the systems

“We are a deployer” as an enterprise-wide answer. Chapter 1 dealt with this one. Cost: provider obligations discovered late, on the one system the firm white-labelled or fine-tuned.

Assuming fine-tuning is always safe customisation

It is sometimes model provision, sometimes substantial modification, sometimes both. Occasionally neither. Cost: Chapter V or Chapter III obligations acquired without any evidence trail, and no way to reconstruct one, because nobody documented the training decisions at the time.

Confusing the model regime with the system regime

Firms building on a compliant GPAI model conclude the compliance is inherited. It is not. The model provider’s file discharges the model provider’s duties, full stop. Cost: a system-level documentation gap at exactly the layer a deployer or an authority examines first.

Overclassification as false prudence

Declaring everything AI and everything high-risk feels conservative. It is not. It dilutes attention, inflates cost, and produces classification records an authority can see were never reasoned. Cost: real money, and weaker files, paradoxically, because a record that classified everything identically proves that no analysis ever happened.

The remedy for all six is the same artefact. A reasoned, dated classification record per system, made under a stated policy, reviewed on defined triggers. Chapter 5 provides the template and the process.

What an auditor will ask for

For this chapter’s material, expect requests for the following. 

  • Your classification policy, with the standing positions on rule-based systems, statistical models, optimisation software and embedded components, reasons included. 
  • Per system, the definitional analysis: which elements of Article 3(1) are met, on what facts. 
  • For systems built on general-purpose models, the analysis separating model duties from system duties, plus any fine-tuning assessment against the Commission’s compute indicators. 
  • For anything configured, retrained, fine-tuned, rebranded or repurposed, the Article 25 record, concluding that provider status transferred or that it did not, with the facts and the documentation perimeter you relied on. 
  • For white-label arrangements, the contractual allocation of provider duties.

A reader holding those artefacts has answered the question this chapter opened with. Whether the Act applies, and to whom. The next question is when, and Chapter 3 takes up the calendar.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 3: The Timeline as It Actually Stands

No part of the AI Act has generated more confusion than its calendar, and the confusion is recent. For two years the timeline was stable and widely published. Then the Digital Omnibus arrived. It moved some dates, held others, added new ones, and did all of this close enough to the original deadlines that half the commentary in circulation is now simply wrong.

This chapter states the position as at July 2026. It separates what is law from what is imminent, and turns the result into a working compliance calendar. It closes with the argument the deferral headlines buried: the obligations that moved are unchanged in content, and the evidence they require ages in your favour only if you start generating it now.

The original schedule

The Act entered into force on 1 August 2024, with its application staged in four steps [1, Art 113].

From 2 February 2025, the prohibitions in Article 5 and the literacy obligation in Article 4. From 2 August 2025, the general-purpose AI model regime, plus governance provisions and penalties. From 2 August 2026, the general application date, which was going to carry the high-risk regime for Annex III systems and the Article 50 transparency duties. And from 2 August 2027, the high-risk regime for AI embedded in Annex I regulated products.

That was the plan. By late 2025 the supporting infrastructure was visibly behind it. Harmonised standards were incomplete. Many Member States had not designated their authorities. Notified body capacity, for practical purposes, did not exist. So on 19 November 2025 the Commission responded with the Digital Omnibus on AI, a package of targeted amendments [2]. One trilogue failed on 28 April 2026. Political agreement came on 7 May. 

The Parliament endorsed the text on 16 June 2026, the Council gave its final approval on 29 June, and the amending regulation was published in the Official Journal on 24 July 2026 as Regulation (EU) 2026/1744, entering into force on 27 July [1].

What the Omnibus changed, and what it did not

The amendments are targeted. The risk-based architecture survives. The actor model survives. The content of the obligations survives. What moved is timing, plus a short list of substantive adjustments.

Deferred

Obligations for stand-alone high-risk systems under Annex III move from 2 August 2026 to 2 December 2027. Sixteen months. Obligations for high-risk AI inside Annex I regulated products move from 2 August 2027 to 2 August 2028. The Member State duty to set up at least one regulatory sandbox moves to 2 August 2027. And the machine-readable marking duty for generative systems under Article 50(2) moves to 2 December 2026 for systems already on the market, a four-month reprieve, not the six first proposed.

Unmoved

Everything already in application stays in application. The prohibitions. Article 4. The GPAI model regime. The rest of the Article 50 transparency duties, telling people they are talking to an AI, labelling deepfakes, disclosing emotion recognition, arrive on 2 August 2026 exactly as originally scheduled. So do the enforcement powers arriving that day [8], including the Commission’s power to fine GPAI model providers. The obligation and the fine switch on together.

Added

From 2 December 2026, Article 5 prohibits AI systems for generating or manipulating non-consensual intimate material and child sexual abuse material [8]. The prohibition covers placing such systems on the market, marketing them without adequate safeguards against such use, and deploying them for that purpose.

Adjusted

The Machinery Regulation moved from Section A to Section B of Annex I. The safety component definition narrowed. Small mid-cap enterprises gained proportionality measures. The AI Office’s supervisory role was strengthened for systems built on a provider’s own GPAI model. And a pathway opened for using special categories of personal data in bias detection. Each of these surfaces in the later chapter where it operates.

The consolidated calendar

Here are the working dates, all now law, verified against the published text of Regulation (EU) 2026/1744:

DateWhat appliesWho is addressed
2 February 2025Article 5 prohibitions; Article 4 AI literacyEveryone in scope
2 August 2025GPAI model obligations (Articles 53 to 55); governance; penalties frameworkGPAI model providers
2 August 2026Article 50 transparency (interaction disclosure, deepfake and emotion recognition disclosure); GPAI enforcement powers and fines; market surveillance structuresProviders and deployers of in-scope systems; GPAI model providers
2 December 2026Article 50(2) machine-readable marking for systems on the market before 2 August 2026; new Article 5 prohibitions on non-consensual intimate material and CSAMProviders of generative systems; everyone, for the prohibitions
2 August 2027Member State regulatory sandboxes; compliance date for GPAI models placed on the market before 2 August 2025Member States; legacy GPAI model providers
2 December 2027High-risk obligations, Annex III stand-alone systemsProviders, deployers, importers, distributors of Annex III systems
2 August 2028High-risk obligations, Annex I embedded systems; delegated acts under the revised Machinery Regulation approachProviders of AI in regulated products; the Commission

Two footnotes to the table matter in practice. First, systems placed on the market after the relevant dates comply from day one of being on the market. The December 2026 marking date is a grandfathering rule for existing systems, not a holiday for new ones. 

If you launch a generative product this autumn, the reprieve belongs to your older competitors, not to you. Second, GPAI models that were already on the market before 2 August 2025 get until 2 August 2027 to bring their documentation into line. 

Which is why some model providers’ files look thinner than the regime suggests they should [9]. It is an explanation. It is not a comfort.

Building the compliance calendar

A regulatory timeline is not yet a compliance calendar. The dates above tell you when obligations bite. Your calendar has to run backwards from each date, through all the work that date quietly assumes you have done. Four steps.

Step one, filter the timeline through your inventory

Take the system list and the role determinations from Chapters 1 and 2, and mark which rows of the table actually address you. A firm with no Annex III systems and no generative products has a short calendar. A firm with one white-labelled chatbot has a longer calendar than it thinks it does.

Step two, break each date into lead-time items

Article 50 disclosure due in August 2026 breaks down into interface changes, wording approval, vendor coordination and staff instructions, and the critical path usually runs through the vendor, who is on nobody’s calendar but their own. 

High-risk readiness for December 2027 breaks down into the entire Part II and Part III programme of this book. The longest items, a working risk management system and a representative testing record, take twelve to eighteen months to build properly. 

And if your conformity assessment needs a notified body, add the body’s queue. You do not control the queue.

Step three, add the trigger-based reviews

Some calendar entries are events, not dates. A system gets modified. A vendor gets swapped. A new use case appears. An Article 25 analysis falls due. Name the triggers, and name the person who watches for each one.

Step four, assign owners and evidence

Every entry carries a named owner and names the artefact that proves completion. A calendar without evidence columns is a to-do list. A calendar with them is the skeleton of your audit file.

One discipline governs all four steps. The Omnibus is now in the Official Journal, so the dates above are settled law, but the lesson of the transition stands. 

A calendar that records its own assumptions, and carries a review entry that fires when the law moves, survives contact with legislative reality. One that silently absorbs headlines does not, and the law will move again.

Why deferral is not a holiday

The deferral changed exactly one variable. The obligations due in December 2027 are, in substance, the obligations that were due in August 2026. The Commission’s stated reason for delaying was that the compliance infrastructure was not ready [5]. Not that the requirements were too demanding. Nothing was removed from the file you must eventually produce.

Three properties of that file reward starting now.

The first is duration evidence. Article 9 requires a risk management system that runs continuously, through the whole lifecycle. A process cannot be evidenced retrospectively. Its evidence is the record of it having run. An organisation that starts in mid-2027 can show an authority a risk management system three months old on the application date. One that kept going from 2026 shows eighteen months of review cycles, updates and closed findings. The obligations are identical. The files are not comparable.

The second is timestamp credibility. Chapter 17 returns to this at length, but the short version belongs here. Classification decisions, data governance records and oversight designs all carry dates. A file where every document was created in the quarter before the deadline reads as reconstruction. Authorities read files this way. So do acquirers and enterprise customers. They read them this way because it is the correct way to read them.

The third is dependency lead time. The deferral bought time partly for the standards bodies, and harmonised standards will land through 2026 and 2027. An organisation with a running programme absorbs each standard as it arrives, adjusting an existing system. An organisation that waited has to implement the standards and build the system simultaneously, in a market where every competitor is hiring the same scarce compliance people in the same quarter.

There is an honest counterargument, and it deserves stating rather than dodging. Work done early risks rework when standards land. Money spent on compliance for a product that may pivot is money at risk. 

Both points are true. Both are managed by sequencing, not by delay. The artefacts least likely to be invalidated by future standards are precisely the foundational ones: 

  • The inventory. 
  • The classifications. 
  • The role determinations. 
  • The data provenance records. 
  • The logging design. 

Build those now. Leave the standard-dependent layers, testing protocols, conformity assessment mechanics, on the calendar for when their dependencies resolve. Chapter 21 sequences the whole thing.

The pattern to avoid has a name in every compliance discipline: the stood-down programme. It is already visible in the market. Organisations read “deadline delayed”, reallocated the budget, dispersed the team. 

In late 2027 they will rehire, rebuild, and try to backfill eighteen months of evidence that cannot be backfilled. The deferral was designed as a runway. A runway only helps an aircraft that keeps moving.

What an auditor will ask for

For this chapter’s material, expect requests for the following: 

  • The compliance calendar itself, with obligations mapped to systems, owners, lead-time breakdowns and evidence columns. 
  • The version history of that calendar, showing it was maintained through the Omnibus transition rather than invented after it. 
  • The record of your Omnibus impact assessment, which of your dates moved, which held, and who signed off the revised plan. 
  • For any programme that paused in 2026, the decision record explaining the pause and the restart plan. 

The gap will be visible in your evidence timestamps whether you explain it or not. Better to be the one explaining.

The calendar tells you when. The next chapter turns to the obligations that have applied longest and carry the highest fines. Get ready for the prohibitions.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 4: Prohibited Practices

Everything else in the AI Act is a compliance programme. Article 5 is a wall.

The practices it lists cannot be cured by documentation, oversight or assessment. The only compliant response to a prohibited practice is not to do it. The only remediation is to stop. These prohibitions have applied since 2 February 2025, longer than any other operative provision, and they sit at the very top of the penalty scale: fines of up to 35 million euros or 7 per cent of worldwide annual turnover, whichever is higher [1, Arts 5, 99(3)].

Most organisations respond to Article 5 with a reflex. We obviously do none of this. The reflex is usually correct. It is also always insufficient. Several of the prohibitions have edges that run straight through ordinary commercial products, recommendation engines, workplace analytics, customer scoring, and the question an authority or a counterparty will eventually ask is not whether you believe you are clear of them. It is whether you checked.

This chapter covers what the prohibitions say, where their edges lie, and how to run and record the check.

The eight prohibitions

Article 5(1) lists eight practices [1, Art 5(1)–(7)]. Here they are, compactly, with their operative limits.

Subliminal and manipulative techniques, point (a)

Placing on the market or using an AI system that deploys subliminal techniques beyond a person’s consciousness, or purposefully manipulative or deceptive techniques, with the object or effect of materially distorting behaviour and causing, or being likely to cause, significant harm. Every element matters. The technique, the material distortion, the significant harm. Ordinary persuasion is not caught.

Exploitation of vulnerabilities, point (b)

Same structure, distortion plus harm, but the lever is exploiting vulnerabilities that come from age, disability, or a specific social or economic situation.

Social scoring, point (c)

Evaluating or classifying people over time based on social behaviour or personal characteristics, where the resulting score leads to detrimental treatment in contexts unrelated to where the data came from, or treatment out of all proportion to the behaviour. Note who is covered. Not just governments. Private actors are inside this one too.

Predictive criminal risk assessment, point (d)

Predicting whether someone will commit a crime based solely on profiling or on their personality traits and characteristics. Systems that support a human assessment already grounded in objective, verifiable facts directly linked to criminal activity fall outside.

Untargeted facial image scraping, point (e)

Building or expanding facial recognition databases by scraping faces from the internet or from CCTV, untargeted. Notice something about this one. There is no distortion test, no harm test. The practice itself is banned.

Emotion recognition in workplaces and education, point (f)

Inferring people’s emotions at work or in educational institutions, except for medical or safety reasons. The exception is narrow. And the setting-based structure means the very same system can be lawful in one deployment and prohibited in the next.

Biometric categorisation for sensitive attributes, point (g)

Sorting people by biometric data to deduce race, political opinions, trade union membership, religious or philosophical beliefs, sex life or sexual orientation. Labelling or filtering lawfully acquired biometric datasets sits outside, as does law enforcement categorisation of that kind.

Real-time remote biometric identification for law enforcement, point (h)

In publicly accessible spaces, prohibited, subject to a short, exhaustive list of exceptions. Targeted search for victims. Prevention of specific imminent threats. Locating suspects of listed serious offences. Even the exceptions come wrapped in prior authorisation, fundamental rights impact assessment and registration requirements.

The Commission published guidelines on the prohibitions on 4 February 2025 [11]. Non-binding, but detailed, with worked examples for each practice. They carry the same weight and the same caveat as the definition guidelines from Chapter 2. Cite them. Do not outsource your reasoning to them. 

And remember that interpretation ultimately belongs to the Court of Justice, not to anyone’s guidance.

The ninth and tenth: December 2026

The Digital Omnibus adds two prohibitions, effective 2 December 2026.

The first covers AI systems that generate or manipulate realistic depictions of an identifiable person’s intimate parts, or of an identifiable person engaged in sexually explicit activity, without that person’s freely given, specific, informed, unambiguous and explicit consent [4]. The second covers systems generating or manipulating child sexual abuse material within the meaning of Directive 2011/93/EU.

Now look at the structure, because it reaches further than the headline suggests. The new prohibitions catch placing a system on the market for these purposes, yes. But they also catch placing a generative system on the market without reasonable safety measures to prevent such generation [7]. And deploying a system for these purposes.

Which means a provider of an image or video generation tool who never intended any abuse must still evidence its safeguards, assessed against the state of the art, by December 2026. If your inventory contains generative media capability, this is a live calendar item today, and the misuse risk assessment belongs in your 2026 workplan. Not your 2027 one.

The state of enforcement

Candour requires a plain statement here. As at the time of writing, no public enforcement action under Article 5 has been announced anywhere in the Union [12]. This despite the prohibitions applying since February 2025 and penalties being available since August 2025.

The gap has structural causes. Many Member States were slow to designate their market surveillance authorities. Most have spread prohibition enforcement across several sectoral regulators, an approach visible in Ireland’s implementing legislation, which assigns different Article 5 practices to the Central Bank, the Workplace Relations Commission, the media regulator and the Data Protection Commission [12]. Four regulators for one article. Coordination takes time.

But drawing comfort from the empty docket would be a mistake, for three reasons.

First, adjacent enforcement already exists. The conduct behind point (e), scraping faces off the internet, is exactly what European data protection authorities fined Clearview AI for, repeatedly, under the GDPR. The authorities that brought those cases now hold AI Act competence, or sit right next to it. A practice can violate Article 5 and the GDPR at the same time, and the GDPR machine has a decade of momentum.

Second, complaints are open to anyone. An affected person. A competitor. An NGO. Enforcement does not have to wait for a regulator’s initiative, and the first wave of complaints is the likeliest trigger for the first wave of decisions.

Third, this is where reputational and regulatory exposure converge. The first Article 5 respondent will be a news story before it is a case number. No general counsel wants that particular distinction.

The empty docket has one practical consequence for your files. There is no enforcement practice to calibrate against yet, so the Commission guidelines and your own documented reasoning carry the full weight. That is an argument for better records. Not thinner ones.

The grey zones

Three areas generate most of the genuine analytical work. Each deserves a standing position in your classification policy.

Emotion recognition

To work out whether a system falls under this prohibition, you need to answer three questions in order.

Question one: is it actually emotion recognition, in the Act’s sense? The legal definition runs through biometric data. That means inferring emotions from a person’s face, voice or physiological signals is covered. Analysing the words someone typed is generally not. So a tool that reads frustration in a customer’s angry email sits outside the prohibition. A tool that reads frustration in the customer’s voice sits inside the concept, and moves to question two.

Question two: where is it used? The prohibition only applies in two settings, workplaces and educational institutions. This produces a result that surprises people. The same voice analytics is prohibited when pointed at your call centre agents, and merely regulated when pointed at the customers they are talking to. One system, one call, two legal outcomes, depending on which end of the line it analyses.

Question three: does the exception apply? The Act permits emotion recognition in those settings for medical or safety reasons. This is a narrow gate. Fatigue detection for lorry drivers passes through it. Genuine clinical applications pass through it. A wellbeing dashboard does not, and neither does engagement scoring wearing safety vocabulary as a costume. If the honest purpose is monitoring productivity or mood, calling it safety does not make it safety.

Now the commercial trap, because it catches organisations that never set out to recognise anyone’s emotions. 

Workplace analytics tools are sold as productivity or wellness software, and somewhere in the vendor documentation sits a quiet mention of emotion, mood or engagement inferred from voice or video. The buyer never asked for emotion recognition. The buyer may now be running a prohibited practice anyway, because the prohibition catches deployers, not just the vendors who built the thing. And workplace emotion recognition is reportedly where the first investigations are looking.

The defence is boring and effective. Screen procurement documents for those words, emotion, mood, engagement, sentiment, before signing. Not after.

Scoring-adjacent systems

Point (c) does not prohibit scoring. It prohibits a specific pathology of scoring. Credit scoring on financial data, insurance pricing on risk data, fraud scoring on transaction data, all lawful, and mostly reappearing as high-risk uses under Annex III rather than as prohibitions. 

The line gets crossed when data from one context drives detriment in an unrelated context, or when the treatment is wildly out of proportion to the behaviour. 

So the screen question for any scoring system is about provenance and consequences. 

  • What feeds the score?
  • What does the score gate?
  • Are those two things related?

A tenant screening product ingesting social media behaviour, or an insurer adjusting premiums on lifestyle data unrelated to the insured risk, sits squarely in the zone where the analysis must be written down.

Manipulation claims

Point (a) frightens marketing departments more than it should, and product teams less than it should. The Commission’s guidelines say plainly that personalised advertising is not inherently manipulative, and the stacked requirements, covert or deceptive technique, material distortion, significant harm, keep ordinary persuasion and personalisation outside the prohibition. 

The zone that deserves real analysis is elsewhere. Engagement optimisation aimed at vulnerable groups. Dark patterns operationalised by AI. Systems that exploit a person’s inferred emotional state to push decisions they would not otherwise take.

The honest screen question is simple to ask and uncomfortable to answer. Does the system’s effectiveness depend on the user not perceiving what it is doing? And could you describe the technique publicly without embarrassment? 

Where the answer to the first is yes, write the analysis and involve counsel before launch. Not after.

Running the prohibition screen

The screen is a short, repeatable assessment applied to every system in the inventory. Its output is a dated record. Its value is that it exists before anyone asks. Five steps.

One, confirm scope. The system is an AI system per your Chapter 2 analysis, and your role in respect of it is recorded, because the prohibitions catch providers and deployers alike.

Two, run the checklist. For each of the ten practices, answer yes, no, or requires analysis, against the system’s actual functionality and its real deployment context. Not its marketing description. For procured systems, the vendor’s documentation and your own configuration both count. The vendor’s compliance statement is an input. It is not an answer.

Three, analyse the flagged items. For anything marked requires-analysis, write a short memorandum applying the elements of the relevant prohibition to the facts, citing the Commission guidelines where they help, and reaching a conclusion. One to three pages is normal. Ten identical pages of boilerplate across ten systems is a signal that no analysis happened at all, and examiners recognise the signal.

Four, decide and record. Clear, prohibited, or conditionally clear. Conditionally clear means lawful in this deployment and prohibited in the one next door, the emotion recognition pattern, and the condition goes into the system’s record as a deployment constraint, with an owner.

Five, set the review triggers. New use case. New population. New data source. Vendor feature update. And the December 2026 additions for anything generative. The screen re-runs on trigger, not on anniversary.

For most organisations the first full screen is a fortnight of work, and the steady state after that is negligible. Weigh the asymmetry. The screen costs days. A prohibition allegation costs the top of the fine scale, mandatory cessation of the practice, and the public record of both.

What an auditor will ask for

For this chapter’s material, expect requests for the following: 

  • The prohibition screen record for every system in the inventory, dated, owned, versioned. 
  • The analysis memoranda for every flagged system, showing the elements of the relevant prohibition applied to actual facts. 
  • For workplace or education deployments with any affect-related functionality, the emotion recognition analysis and its deployment constraints. 
  • For scoring systems, the data provenance and consequence map. 
  • For generative media systems, the misuse safeguards assessment against the December 2026 prohibitions, plus evidence the safeguards actually operate. 
  • For anything conditionally clear, the record showing the condition reached the teams capable of breaching it.

The prohibitions are the Act’s floor. The next chapter climbs to the decision that determines everything above the floor. Whether a system is high-risk.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Part II — Risk Governance

Part I told you what the Act is and when it bites. Part II is where the work starts.

The five chapters ahead cover the decisions and structures that everything else depends on. Which of your systems are high-risk, and how to prove you checked. The risk management process that must run for as long as the system does. The data governance that must be in place while the data work happens, not after. The people and committees who own all of it. And the literacy your staff must actually have, not merely have been emailed about.

A warning before you start. This is the part of the book where reading stops being enough. Every chapter in Part II ends with a list of records an auditor will request, and none of those records writes itself. If Part I was the map, Part II is the part where you dig.

Chapter 5: Classification, the Load-Bearing Decision

Every obligation in Chapter III of the Act hangs from a single determination. Is the system high-risk?

Get it right, and the rest of this book is a build plan. Get it wrong in one direction, and you construct a conformity assessment nobody required. Get it wrong in the other, and you discover the requirement from a market surveillance authority, with the application date behind you and the evidence unbuilt.

Classification is also the determination most often done badly. Not because it is hard. Because it is done casually. It gets treated as a question with an obvious answer, decided in a meeting, recorded nowhere, never revisited. 

This chapter sets out the legal tests first, then the discipline: classification as a repeatable, versioned process producing one artefact per system. Including the artefact almost everyone forgets. The documented negative.

The two routes to high-risk

Article 6 creates two independent routes. A system is high-risk if it travels either one.

The first route, Article 6(1), runs through product law. A system is high-risk where two conditions meet. It is a product, or a safety component of a product, covered by the Union harmonisation legislation listed in Annex I. And that product must undergo third-party conformity assessment under that legislation [1, Art 6(1)].

Annex I lists the familiar product regimes. Machinery, toys, lifts, medical devices, in vitro diagnostics, radio equipment, vehicles, aviation, and more. 

The logic is inheritance. Where Europe already treats a product category as dangerous enough to need third-party checking, AI doing safety work inside that product inherits the concern.

Two Omnibus amendments reshaped this route. The safety component definition was narrowed. Systems used solely for user assistance, performance optimisation, efficiency, automation, convenience or quality control no longer qualify, unless their failure could endanger health or safety [4]. 

The Machinery Regulation moved from Section A to Section B of Annex I, which takes AI-enabled machinery out of the dual-compliance model. AI-specific requirements for machinery will arrive instead through delegated acts, by August 2028.

For manufacturers, the practical effect is that the Annex I analysis now starts with a failure question. What does the AI component do, and if it malfunctions, does anyone get hurt?

An optimisation layer whose worst failure is inefficiency is out. A perception layer whose worst failure is a collision is in. And one further carve-out arrived in the final text: a product required to undergo third-party assessment solely for risks unrelated to health and safety, radio spectrum allocation or electromagnetic interference being the named examples, does not satisfy the Article 6(1) condition at all [1, Art 6(1c)].

The second route, Article 6(2), runs through use cases. A system is high-risk where it falls within an area listed in Annex III [1, Annex III]. The Annex lists eight:

Biometrics, meaning remote identification, categorisation by sensitive attributes, and emotion recognition where it is not already prohibited. 

  • Safety components of critical infrastructure. 
  • Education and vocational training, admission, assessment, proctoring. 
  • Employment and worker management, recruitment, selection, promotion, termination, task allocation, monitoring. 
  • Access to essential services, creditworthiness, life and health insurance pricing, emergency dispatch, benefit eligibility. 
  • Law enforcement. 
  • Migration, asylum and border control. 
  • The administration of justice and democratic processes.

Read the Annex itself, not summaries of it. The entries are drawn narrowly and the drafting does real work. 

Take employment. The Annex does not catch every system an HR department uses. It catches systems intended for recruitment or selection, for decisions affecting the terms of work relationships, and for monitoring and evaluating people in them. A payroll calculator is not in the Annex. A CV-ranking model is its first example.

One principle governs both routes. Intended purpose. Classification attaches to the use the provider intends, as documented, marketed and instructed. 

Which is why this chapter and Chapter 2’s Article 25 analysis interlock. A deployer who repurposes a general-purpose system into an Annex III use has changed the classification, and quite possibly changed who the provider is.

The Article 6(3) filter

The Annex III route has an escape valve, and this is where most classification work concentrates. Under Article 6(3), a system that falls within an Annex III area is nevertheless not high-risk where it does not pose a significant risk of harm to health, safety or fundamental rights, including by not materially influencing the outcome of decision making [1, Art 6(3)].

The provision then lists the conditions under which that is the case. A system qualifies where it is intended to perform a narrow procedural task. Or to improve the result of a previously completed human activity. Or to detect decision-making patterns, or deviations from prior patterns, without replacing or influencing the human assessment. Or to perform a task preparatory to an assessment relevant to the Annex III use case.

And one counter-rule overrides all of it. A system that performs profiling of natural persons is always high-risk, whatever else is true of it [1, Art 6(3)].

Applied honestly, the filter separates infrastructure from influence. A system that translates, transcribes, deduplicates, formats, routes or indexes inside a recruitment or credit process is doing narrow procedural or preparatory work. 

A system that scores, ranks, shortlists, flags or recommends is influencing the outcome, and the filter does not save it.

The self-deception to guard against here is the word “recommendation”. Teams argue that their tool only recommends, a human decides, so the influence is not material. The logic of Recital 53 and the structure of the filter run the other way. A ranking the human works from is the textbook case of material influence on an outcome. The human in the loop is what Article 14 regulates. It is not what Article 6(3) exempts.

The filter carries its own paperwork. A provider concluding that its Annex III-area system is not high-risk must document that assessment before placing the system on the market, and must hand it to national competent authorities on request [1, Art 6(4)]. Registration in the EU database survived the Omnibus too, though in simplified form. 

The Commission had proposed deleting the registration duty for filter-exempted systems. Parliament and Council both refused, and the final text keeps registration while trimming what goes into the public entry. 

Notably, the short summary of the exemption grounds was removed from the register. Which means your internal Article 6(4) assessment is now the only place your reasoning lives. That makes the private record more load-bearing, not less. 

One timing note so nobody panics: the duty bites when Chapter III does, 2 December 2027 for Annex III systems. The duty survived. It is not imminent.

Note the asymmetry of risk around the filter. The Commission holds the power to issue guidelines and to amend the filter conditions. And an authority reviewing your Article 6(3) conclusion reviews it with hindsight, after whatever incident brought them to your door. 

The filter is legitimate. Using it is not aggressive but using it without a written assessment is.

The current state 

The Commission has now used that power, in part. In May 2026 it published draft guidelines on Article 6 classification, in three documents: general principles, the Annex I route, and the Annex III route [30]. Consultation closed in July and the final text is expected by the end of 2026, so what follows is the Commission’s current thinking rather than settled guidance.

Three points in the draft matter for the way this chapter has told you to work.

Intended purpose

The draft treats intended purpose as determinative and which must be described consistently across the instructions for use, the technical documentation and the marketing. A mismatch between what a system does in practice and how its purpose is described will not protect a provider from classification. That is the same reconciliation Chapter 10 asks you to run, now with the Commission saying out loud what it will look for.

Human oversight

The second confirms the position this chapter took on the filter. Human oversight does not affect classification. A provider cannot avoid high-risk status by building a human into the loop, because oversight is a compliance requirement for systems already classified as high-risk, not a design feature that prevents classification. If you have been arguing the other way internally, that argument is now harder to make with a straight face.

The third is new and catches composite products. Where multiple AI components interact and their combined outputs materially influence a decision in a high-risk use case, the draft treats them as a single AI system. Organisations that classified component by component, and found each one individually below the line, should re-run the analysis on the assembly.

The draft also reads the Article 6(1) third-party conformity assessment condition more broadly than many manufacturers assumed, which pulls more embedded AI into scope rather than less. Add the final guidelines to the trigger list. When they publish, every filter position you hold gets re-read against them.

Documenting the negative: the artefact everyone forgets

Compliance files fail in a characteristic way. They hold reasoned records for the systems classified as high-risk, and silence for everything else. Authorities, acquirers and enterprise customers all read that silence the same way, as absence of analysis. Usually, that is exactly what it is.

The negative classification record fixes this, and it is the cheapest high-value artefact in the entire programme. For every system in the inventory that is not high-risk, write a short record saying why not.

The reasoning differs by exit point, and the record should show which exit was taken. Outside the AI system definition entirely, Chapter 2 analysis attached. An AI system, but within no Annex III area and no Annex I regime, naming the areas considered nearest and why the system falls outside them. 

Or within an Annex III area but exempt under the filter, stating the condition relied on and addressing the profiling counter-rule. Providers using the filter also complete the formal Article 6(4) assessment.

Three disciplines make the record worth having. Name the nearest miss. 

A record that says “we considered Annex III point 4 on employment, the system allocates desks, not tasks or evaluations, for these reasons” demonstrates analysis in a way “not high-risk” never will. Address profiling expressly wherever the filter is used, because it is the override an examiner checks first. Add the date and sign the whole thing. An undated negative is a shrug.

There is a commercial payoff here too. Procurement questionnaires and due diligence requests ask for your high-risk systems, and increasingly they ask for evidence that the rest were assessed. Chapter 17 comes back to the evidence map. 

The point for now is that a file of documented negatives converts the largest part of your inventory from unexamined exposure into demonstrated diligence. At the cost of a page per system.

Classification as a process, not a memo

A classification is a conclusion about a system at a moment, under the law as it stood, for a purpose as then defined. All of those things move, systems gain features, purposes drift. The Commission can amend Annex III and the filter, and the Omnibus has already proved the framework itself moves. A one-off memo starts decaying the day it is written.

The remedy is to run classification the way you run any other control. A defined process, with inputs, owners, outputs and triggers. Five elements to memorise:

  • A classification policy: one document stating which tests are applied, in what order, by whom, plus the standing positions from Chapters 2 and 4. 
  • What your organisation treats as an AI system. 
  • How it reads the filter conditions. 
  • What it does with vendor claims. 
  • Individual records then cite the policy. 

This keeps a hundred records consistent, and it shows an authority a method rather than a pile of improvisations.

A standard record

One template per decision:

  • System identity and version. 
  • Intended purpose as documented. 
  • Role determination. 
  • The Annex I analysis. 
  • The Annex III analysis with the nearest-miss reasoning. 
  • The filter analysis where used. 
  • The profiling check. 
  • The conclusion, the decider, the date, and the review triggers. Appendix C provides it. 

A completed record should be readable by an outsider in five minutes and defensible in an inspection without its author in the room.

Defined deciders

Classification is a legal judgement made on technical facts. So the process names who supplies the facts, usually the product owner, and who owns the judgement, compliance or counsel, with an escalation route for the genuinely contestable calls. 

A classification nobody owns is a classification nobody defends.

Versioning

Records are never overwritten, they are superseded. When a system changes and gets reclassified, the new record references the old one and states what changed. The chain that results is itself evidence. It shows the authority a living process and also protects you internally, two years later, when a new team wonders why a call was made.

Triggers

Reclassification fires on events, not anniversaries. For example, material feature changes, new purposes or new user populations, fine-tuning or retraining beyond the documented perimeter. It could be vendor substitution or regulatory change, which is not hypothetical. 

The Omnibus narrowed the safety component definition and moved the Machinery Regulation, and that alone re-opens every Annex I classification a manufacturer made before mid-2026. Someone must own the watch on Annex III amendments and filter guidelines. The Commission holds delegated powers over both.

Sized honestly, none of this is a bureaucracy. A mid-sized deployer’s first full classification pass over a fifty-system inventory is a few weeks of work, concentrated on the ten systems near the lines. Steady state is triggered maintenance. 

Now price the alternative, from Chapter 3. A high-risk obligation discovered in 2027 with no evidence trail. Or a filter position taken silently, examined after an incident. The weeks look cheap.

What an auditor will ask for

For this chapter’s material, expect requests for the following. 

  • The classification policy. 
  • The complete set of classification records, positive and negative, reconciling exactly to the inventory. Every inventory entry has a record. Every record has an inventory entry. 
  • For Annex I products, the safety component analysis under the post-Omnibus definition, failure-mode reasoning included. 
  • For each Article 6(3) exemption, the documented assessment in Article 6(4) form, the condition relied on, the profiling check, and the simplified registration entry once the duty bites. 
  • For each high-risk conclusion, the handover into the Chapter III programme that Parts II and III of this book build. 
  • The version chains for anything reclassified, plus the trigger log showing the process actually fired when the Omnibus changed the law.

Classification decides which systems carry the heavy obligations. The next chapter begins the heaviest of them. The risk management system of Article 9.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 6: The Risk Management System

Article 9 is where the AI Act stops describing your system and starts describing your organisation. Every obligation before this point could, at least in principle, be satisfied by examining the system itself. This one cannot. It demands a process that runs for as long as the system does. And it demands proof that the process ran.

The Article opens with a deceptively short sentence. A risk management system shall be established, implemented, documented and maintained in relation to high-risk AI systems [1, Art 9(1)].

Four verbs. Auditors read all four. Established means it exists on paper, with an owner. Implemented means it demonstrably operated. Documented means the operation left records. Maintained means the records show revision over time. A risk assessment done once, however thorough, satisfies exactly one verb out of four.

What “continuous iterative process” actually means

Article 9(2) defines the risk management system as a continuous iterative process, planned and run throughout the entire lifecycle of the system, requiring regular systematic review and updating [1, Art 9(2)]. Every word in that phrase is doing a job, and most compliance programmes underestimate all of them.

Continuous does not mean constant. Nobody expects daily risk reviews. It means the process has no end date and no finished state. Your risk management system for a given system closes on one day only, the day the system is withdrawn.

Iterative means the process revisits its own conclusions. Each cycle takes the previous cycle’s outputs as inputs. A risk estimated as low in version one gets re-estimated in version two, and the record shows both estimates, plus the reason anything changed.

Planned means the cadence was set in advance, not improvised when convenient. Your risk management plan states when reviews happen and which events force an unscheduled one. And the events matter more than the calendar. A substantial modification. A serious incident. Bad news from post-market monitoring. A change in the law. Each of these should trigger a review cycle, and your plan should say so before any of them happens.

Here is the operational test, and it is brutally simple. Pick any date in the system’s life. Ask what the risk management system was doing that month. If the honest answer is nothing, and no plan explains why nothing was scheduled, then the process was not continuous. It was episodic. And the documentation will show it.

The four steps

Article 9(2) prescribes the cycle’s content in four steps. The drafting rewards close reading, because each step has a different risk universe.

Step one is identifying and analysing the known and reasonably foreseeable risks the system can pose to health, safety or fundamental rights, when used as intended [1, Art 9(2)(a)]. Two boundaries matter here. The risk universe is health, safety and fundamental rights. Not commercial risk, not reputational risk. And at this step, the lens is intended use only.

Step two widens the lens. It requires estimating and evaluating the risks that may emerge when the system is used as intended, and under conditions of reasonably foreseeable misuse [1, Art 9(2)(b)]. Misuse is a defined term: use that departs from the intended purpose, but which may result from reasonably foreseeable human behaviour or interaction with other systems [1, Art 3(13)].

The word reasonably carries the load. You are not required to imagine every abuse a creative person could invent. You are required to anticipate the misuses your deployment context makes predictable, and, just as important, to write down which ones you considered and rejected as not reasonably foreseeable. 

Keep the rejected list. It is your defence, years later, when a regulator asks why a misuse that materialised was never assessed. “We considered it and here is why we ruled it out” is an answer. Silence is not.

Step three brings in the field. Risks identified through post-market monitoring under Article 72 must be evaluated within the risk management system [1, Art 9(2)(c)]. This is the clause that makes the process genuinely continuous rather than notionally so. Your monitoring plan and your risk system are not two documents that happen to coexist. The first feeds the second, and the feeding must be visible in the records.

Step four is adopting appropriate and targeted risk management measures to address what steps one to three found [1, Art 9(2)(d)]. Targeted is the operative word. A measure must trace to a risk. Generic controls adopted because they seemed prudent do not discharge this step, however sensible they are, because nothing connects them to the identification work. A control without a parent risk is an orphan, and auditors collect orphans.

Mitigation and the hierarchy of measures

Article 9(5) imposes an order of preference that anyone from product safety will recognise. First, eliminate or reduce the risk through design and development. What cannot be designed out gets mitigation and control measures. And whatever remains after that is addressed by information and training for deployers [1, Art 9(5)].

The hierarchy has teeth. Reach for the instruction manual as your primary control, and you invite the obvious question: why could this not be reduced in design? Your records should show the question was asked in the right order. What did we design out. What did we control. And only then, what did we warn about.

A risk file that jumps straight to warnings tells its own story. It reads like a file assembled after the design was frozen. Which is usually exactly what happened.

Residual risk and the acceptability judgement

After the measures are applied, something is left over. That is residual risk, and Article 9(5) requires the overall residual risk of the system to be judged acceptable [1, Art 9(5)]. This is the least mechanical moment in the entire compliance file. It is also where template thinking fails completely, because no template can make a judgement for you.

Acceptability is a decision, and someone has to make it. The record must show who, on what date, against what criteria, and considering which residual risks, individually and in combination. Define your criteria before making the judgement. Not afterwards, reverse-engineered to bless whatever outcome the project needed. 

A common and defensible approach borrows the ALARP logic from safety engineering: residual risk is acceptable where further reduction would be grossly disproportionate to the benefit gained. 

Whatever standard you choose, adopt it in writing, in advance, and apply it consistently across your whole portfolio. Inconsistency between systems is precisely what a horizontal audit exists to find.

And note who the risk is measured against. Article 9(9) requires specific consideration of whether the system is likely to have an adverse impact on persons under eighteen and other vulnerable groups [1, Art 9(9)]. If your intended purpose or your foreseeable deployment touches such groups, the file must show they were considered separately. Not folded quietly into a general user population.

Testing

Articles 9(6) to (8) make testing part of risk management, not a separate engineering hobby. High-risk systems are tested to identify the most appropriate measures, to ensure consistent performance for the intended purpose, and to confirm compliance with the Chapter III requirements. 

Testing runs against prior defined metrics and probabilistic thresholds appropriate to the purpose, and may include real-world testing under Article 60 [1, Arts 9(6)–(8), 60].

The phrase prior defined is the audit hook. Metrics chosen after the results were known prove nothing. Your file should contain the test plan, with its thresholds, dated before the test records that report against them. This discipline costs nothing at the time. It is impossible to retrofit. That combination should tell you when to do it.

Integrating with ISO 31000 and ISO 42001

If your organisation already runs risk management under ISO 31000, or an AI management system under ISO/IEC 42001, you are not starting from zero. The Act says so itself. Article 9(10) lets providers already subject to risk management requirements under other Union law combine the Article 9 process with their existing procedures, provided the combination achieves an equivalent level of protection [1, Art 9(10)]. And whatever the strict scope of that provision, nothing stops you running one process that serves both masters.

But the mapping is close, not congruent. Three gaps recur, and each is a finding waiting to happen.

First, the risk universe differs. ISO 31000 manages risk to organisational objectives. Article 9 manages risk to the health, safety and fundamental rights of the people your system affects. These overlap far less than would be convenient. 

Your enterprise risk register will faithfully capture the model drift that threatens revenue, and completely miss the same drift’s disparate impact on a protected group, because no organisational objective was harmed. Nobody did anything wrong. The frame just cannot see it. 

The fix is a dedicated risk universe inside your existing process, an Article 9 category whose identification prompts are written in fundamental rights language.

Second, the rhythm differs. ISO frameworks tolerate annual review cycles for stable risks. Article 9’s continuous process, fed by post-market data, expects the cycle to turn whenever the field produces a signal. If your ISO calendar is annual, your Article 9 layer needs event triggers sitting on top of it, written into the integrated procedure.

Third, the acceptability decision differs, and this one is subtle. ISO 31000 lets the organisation set its own risk appetite against its own objectives. Article 9’s acceptability judgement is made against the interests of the people affected, and the Act, not your board, decides whose interests count. An integrated process must keep these two decisions visibly separate. A residual risk your organisation cheerfully accepts on commercial grounds may still fail the Article 9 test. One record must never be dressed up as the other.

ISO/IEC 42001 narrows all three gaps considerably, because its risk and impact assessment machinery was drafted with this Regulation in view. An organisation with a working 42001 system typically needs to extend its impact prompts to the full fundamental rights range, wire post-market monitoring into the review triggers, and formalise the residual risk acceptance record. 

Real work, but weeks of it, not quarters. And the certification artefacts double as Article 9 evidence with modest relabelling.

What 42001 certification does not do is discharge Article 9 by itself. Certification proves a management system exists. Article 9 wants proof that the system processed this AI system’s risks, on these dates, with these outcomes. A certificate on the wall answers the first question. Only records answer the second.

Timing

Article 9 sits in Chapter III and follows the deferred high-risk timeline. 2 December 2027 for Annex III systems, 2 August 2028 for Annex I [4].

The deferral changes when the obligation bites. It changes nothing about the arithmetic. A continuous lifecycle process cannot be conjured on its deadline. It has to have been running already, or it has no history to show. 

Providers planning to place Annex III systems on the market in 2027 should treat 2026 as the year the process starts generating records. A risk management system born on its own compliance date arrives with exactly the episodic file this chapter has been warning about.

What the auditor asks for

The evidence set for Article 9 is the longest in the book so far, and every item follows from the text above.

  • The risk management plan, owned and dated, stating cadence and event triggers. 
  • The risk register for the specific system, showing identification, estimation and evaluation across intended use and foreseeable misuse, including the considered-and-rejected misuse list. 
  • The linkage records showing post-market data entering the review cycle. 
  • The measures register, each control traced to an identified risk, ordered by the hierarchy, design first, control second, warnings last. 
  • The residual risk acceptance record, with a named decision-maker, a date, criteria, and reasoning. 
  • The vulnerable groups analysis wherever Article 9(9) is engaged. 
  • The test plan with prior defined metrics, dated before the test records. 
  • The version chain across all of it, because the iteration is the compliance.

A specimen risk register and residual risk acceptance record appear in the appendices. The working versions, with the full prompt sets and the integrated ISO mapping, are part of the template collection described at the back of this book.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 7: Data and Data Governance

Article 10 is the part of the Act most often summarised and least often read. The summary everyone repeats says training data must be relevant, representative and free of errors. The Article actually says something more careful. More demanding in some places, more forgiving in others. And the difference decides what your dataset documentation has to prove.

The structure is simple enough. High-risk AI systems that involve training models with data must be developed on the basis of training, validation and testing data sets that meet the Article’s quality criteria, and, wherever special category data is used, the conditions of Article 4a [1, Art 10(1)].

Two things to notice straight away. The obligation covers all three data set roles, not training data alone. And it attaches through governance practices, not through the data’s intrinsic properties.

That second point is the one to internalise first, because it changes what compliance means here.

Governance before quality

Article 10(2) requires the data sets to be subject to data governance and management practices appropriate for the intended purpose [1, Art 10(2)]. Then it lists what those practices must cover. 

  • The design choices.
  • The collection processes and where the data came from. 
  • The preparation operations, annotation, labelling, cleaning, updating, enrichment, aggregation.
  • The assumptions, especially about what the data are supposed to measure and represent. 
  • An assessment of whether enough suitable data even exists. 
  • Examination for possible biases, and measures to detect, prevent and mitigate them. 
  • The identification of gaps and shortcomings, plus what was done about them.

Now read that list the way an auditor reads it. This is not a quality standard. It is a documentation standard about decisions. Every item on the list is something someone chose or assumed, and the compliance artefact is the record of the choosing. 

Where did the data come from, and why that source? What does this label mean, and who decided? What is the dataset supposed to measure, and what is the argument that it measures it? Which gaps were found, and what happened next?

This leads to a conclusion that surprises people, so here it is plainly. An organisation can hold flawed data and a compliant governance record, because the record honestly documents the flaw, assesses its impact, and shows the mitigation. An organisation cannot hold excellent data and no record. Article 10(2) regulates the practices. Undocumented practices did not happen.

What “relevant, representative, free of errors” actually says

Article 10(3) is where the misquotation matters. The actual text requires that training, validation and testing data sets be relevant, sufficiently representative, and, to the best extent possible, free of errors and complete in view of the intended purpose [1, Art 10(3)].

Three qualifiers do the work. Sufficiently representative. Not representative, full stop. To the best extent possible. Not absolutely. And in view of the intended purpose, which conditions everything that comes before it.

The drafters knew perfect data does not exist, and the Article does not demand it. What it demands is defensible adequacy against a stated purpose. Which turns each adjective into a question your documentation must answer.

Relevant asks: what is the demonstrated relationship between these data and the task the system performs? Your answer is the measurement argument from 10(2), stated for the deployed context, not the laboratory.

Sufficiently representative asks: representative of what population, under which deployment conditions? Article 10(4) sharpens the question. Data sets must take into account, to the extent the intended purpose requires, the characteristics particular to the specific geographical, contextual, behavioural or functional setting where the system will be used [1, Art 10(4)]. 

A recruitment model trained on one labour market and deployed in another fails this test. Not because the data are bad. Because the documentation cannot connect them to the setting. Your record should name the intended settings, and show, setting by setting, why the data suffice or what was done where they did not.

Free of errors, to the best extent possible, asks: what error detection did you run, what error rate did you find, and why is the residual rate tolerable for this purpose? Notice what counts as evidence here. Not a claim of cleanliness. The cleaning log. The known-error register. The tolerance rationale. A dataset documented as containing known errors, with assessed impact, reads as governed. A dataset documented as error-free reads as unexamined. Perfection on paper is a red flag, not an achievement.

Complete asks the mirror question. What should be in the data that is not, and what follows from the absence? This is the gap item from 10(2) restated as a quality property, and one record answers both.

Bias examination and the Article 4a pathway

The bias obligations sit in two places. Article 10(2)(f) and (g) require examination for biases likely to affect health and safety, harm fundamental rights, or lead to discrimination prohibited under Union law. Especially where the system’s outputs influence its future inputs. Plus appropriate measures to detect, prevent and mitigate whatever the examination finds.

That feedback-loop clause deserves its own sentence in your documentation. Where outputs return as future training inputs, yesterday’s skew compounds. The examination must address the loop, not just the snapshot.

Then comes the practical problem every practitioner knows. Testing whether a model discriminates by ethnicity, health status or religion generally requires processing exactly the special categories of personal data the GDPR restricts most tightly. You need the sensitive data to prove you are not misusing the sensitive data. For years this was an awkward circle.

Article 4a is the Act’s answer, and it is new. The Digital Omnibus inserted it as a free-standing article near the front of the Regulation. Before the amendment, the only route through the circle sat inside Article 10 and reached providers of high-risk systems alone. Article 4a is wider on both counts. It reaches deployers as well as providers, and it reaches systems and models outside the high-risk tier.

The rule itself: to the extent strictly necessary to ensure bias detection and correction in relation to high-risk AI systems, providers may exceptionally process special categories of personal data, subject to appropriate safeguards [1, Art 4a(1)].

Then the conditions stack up, and every one of them is load-bearing.

  • The processing is only permitted where bias detection and correction cannot be effectively fulfilled with other data, including synthetic or anonymised data.
  • The special category data must carry technical limitations on re-use, plus state of the art security and privacy-preserving measures, including pseudonymisation.
  • It must be secured, protected and access-controlled, with the access documented and confined to authorised persons under confidentiality obligations.
  • It must not be transmitted, transferred or accessed by other parties.
  • It must be deleted once the bias is corrected, or when the retention period ends, whichever comes first.
  • The GDPR records of processing must state the reasons why processing special category data was strictly necessary, and why other data could not achieve the objective [1, Art 4a(1)(a)–(f)].

The Omnibus history matters here, because your readers will have seen the headlines. The Commission proposed loosening this pathway. Lowering the threshold from strictly necessary to merely necessary, and extending it beyond providers of high-risk systems. Both the Council and the Parliament pushed back on the threshold. The strict necessity standard was reinstated [14, 15].

On scope, though, the extension survived, in disciplined form. Article 4a(2) opens the pathway to deployers of high-risk systems, and to providers and deployers of other AI systems and models, under every condition and safeguard of paragraph 1, where the biases in question are likely to affect health and safety, harm fundamental rights or produce prohibited discrimination, especially where outputs influence future inputs [1, Art 4a(2)]. The final text adds one clarification in terms: the provision creates no obligation to conduct bias detection at all. It is a door, not a duty.

Two practical consequences follow.

First, timing. Article 4a is not one of the provisions the Omnibus deferred. The deferral moved the Chapter III high-risk obligations to December 2027 and August 2028. The legal basis for bias testing is available before those dates, which matters because the evidence it produces is exactly what those obligations will ask for. Confirm the position against the current text before you rely on it, but plan on the basis that the door is open now.

Second, the deployer extension changes the vendor conversation. A deployer who suspects a procured system of skewed outcomes no longer has to argue that only the provider may lawfully test for it. The testing may now be yours to run, under the full safeguard stack, and the finding feeds Chapter 8’s monitoring duties.

The practical reading: the door exists, the threshold is unchanged, and walking through it still requires the documented necessity analysis, the exhausted-alternatives record, and the full safeguard stack. Cite Article 4a in your own records. Where older material cites the pre-Omnibus position, check it against the current consolidated text rather than assuming either way.

The first document a supervisory authority will request is your analysis of why synthetic or anonymised data could not do the job. Write it before processing begins, not after.

One boundary clarification, because confusion here is common. Article 4a is a permission inside the AI Act, aimed at bias correction. It does not turn your bias testing into a GDPR-free zone. The processing still needs its GDPR Article 9 condition, its records entry, and in most configurations a data protection impact assessment. Your DPO should hold the necessity analysis alongside the AI file. Same facts, two regimes, one story.

Dataset documentation an auditor can actually read

Annex IV requires the technical documentation to describe, among much else, the data sets used. Their provenance, scope and main characteristics. How the data were obtained and selected. Labelling procedures. Cleaning methodologies [1, Annex IV, para 2(d)].

That is the formal hook. The practical question is what a readable dataset record looks like, and the honest answer is that industry converged on it before the Act did. The datasheet. One structured document per data set, answering a fixed set of questions.

A datasheet that discharges Article 10 contains, at minimum, twelve entries:

  • Identity and version, because data sets change and the records must chain. 
  • Role, training, validation or testing, since the 10(3) criteria apply per role and the split rationale belongs in the record. 
  • Purpose statement, what these data are supposed to measure and represent, in 10(2)’s own formulation. 
  • Provenance, sources, collection process, legal basis, licences. 
  • Composition, size, fields, populations covered, time span. 
  • Preparation history, every annotation, labelling, cleaning and enrichment operation, with who did it and when. 
  • The representativeness analysis, tied to the named deployment settings of 10(4). 
  • Known errors and completeness gaps, with impact assessment and the tolerance rationale. 
  • The bias examination, methods, findings across the protected characteristics your context makes salient, and the feedback-loop analysis. 
  • The mitigations, each traced to a finding. 
  • The Article 4a record where the pathway was used, or one line saying it was not needed and why.
  • A sign-off: a named owner and a date. An unowned datasheet is a draft.

The test of readability is not length. It is this. A stranger with regulatory training picks up the record and answers, without asking anyone, the five questions every audit converges on. Where did the data come from? What do they claim to represent? What is wrong with them? What was done about it? 

And who decided that was enough.

Timing and scope notes

Article 10 follows the deferred Chapter III timeline. 2 December 2027 for Annex III systems, 2 August 2028 for Annex I. The arithmetic from the last chapter applies here with extra force, because data governance is retrospective by nature. 

A model trained in 2026 on undocumented data cannot have its collection decisions honestly re-documented in 2027. The people are gone, the decisions are forgotten, the trail is cold. 

The datasheet discipline has to be running while the data work happens. There is no other time it can run.

Note the scope gate in 10(1). The Article’s data set machinery attaches to systems that involve training models. For high-risk systems built without trained models, Article 10(6) applies the governance requirements, and Article 4a(1) where engaged, to testing data sets only [1, Art 10(6)].

This is an often-missed narrowing, worth checking against your portfolio before building datasheets nobody required.

What the auditor asks for

The evidence set:

  • The data governance procedure, owned and dated, covering the full 10(2) practice list.
  • One datasheet per data set per version, in the structure above, with the training, validation and testing split rationale. 
  • The representativeness analysis is tied to named deployment settings. 
  • The known-error and gap registers, with the tolerance decisions. 
  • The bias examination records, methods, findings, mitigations, including the feedback-loop analysis wherever outputs feed future inputs. 
  • The Article 4a file where the pathway was engaged: necessity analysis, alternatives assessment, safeguard implementation, access log, deletion record, and the matching GDPR records entry. 
  • The version chain, because data sets evolve, and the governance must be shown to have followed them.

The specimen datasheet appears in the appendices. The working version, with the full question set, the bias examination prompt library and the 4a(1) necessity analysis template, is part of the template collection described at the back of this book.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 8: Governance Structures and Accountability

The Act never uses the phrase “AI governance officer”. It mandates no role, no committee, no reporting line. This surprises organisations trained by the GDPR, which conjured the data protection officer into thousands of job descriptions across Europe. The AI Act took the opposite approach. It prescribes outcomes and evidence, and leaves the organisational chart entirely to you.

That freedom is a trap for the unprepared. When no role is mandated, accountability defaults to nobody. And nobody produces no records. 

Every obligation in this book so far, the risk cycles of Chapter 6, the datasheets of Chapter 7, the classification records of Chapter 5, quietly presumes that someone exists whose job is to make them happen. This chapter is about constructing that someone.

Who owns AI compliance internally

The honest answer, in most organisations today, is that AI compliance is owned by whoever noticed it first. Usually the DPO. Sometimes legal. Occasionally the CISO. And in engineering-led companies, a product manager who read the Act on a train.

Each default fails in its own predictable way. The DPO reads every AI problem as a data problem, and misses safety and fundamental rights. Legal reads the obligations accurately and cannot verify a single technical claim engineering makes. The CISO governs the models nobody bought through procurement, and misses the ones the business units did buy. The product manager knows the systems intimately and has no authority over anyone else’s.

The structure that survives contact with an audit has three layers. Attach whatever titles you like.

First, a named accountable owner. Senior enough to commit budget and to stop a deployment. This matters, so let me put it bluntly. An owner who cannot say no is not an owner. That is a coordinator.

Second, a cross-functional working level. Legal, data protection, security, engineering, and the business units actually deploying the systems. Why all of them? Because every artefact in this book requires at least two of those disciplines to produce honestly.

Third, named system-level owners. One per entry in the inventory. Portfolio-level accountability without system-level ownership produces a familiar result: beautiful policies and empty risk files.

Size the machinery to the portfolio. An organisation running three vendor chatbots needs an owner, a register and a review rhythm. It does not need a committee. An organisation building Annex III systems needs the full structure, staffed before the December 2027 obligations bite, for the reason Chapters 6 and 7 hammered home. Evidence is generated by processes running in time. And processes need owners before they can run.

One duty already sits on whoever you appoint. Article 4’s literacy obligation applies now, and after the Omnibus it requires providers and deployers to take measures to support the development of AI literacy among staff and anyone else operating AI on their behalf [1, Art 4] [4]. The governance owner’s first deliverable is usually the literacy programme. It is the one obligation with no deferral attached to it.

Board reporting

Nothing in the Act requires board reporting on AI. Do it anyway, for two reasons.

The first is penalty exposure. The fine ceilings are set against worldwide turnover, and worldwide turnover is the board’s risk universe by definition. The second is that several of this book’s artefacts need decisions only senior governance can honestly make. 

Chapter 6’s residual risk acceptance is the clearest case. Someone accepts risk on the organisation’s behalf, and that acceptance record gains real weight when the person accepting demonstrably has the authority to bind the organisation. A junior analyst accepting residual risk on a credit model binds nobody and convinces nobody.

A board pack that serves the file as well as the boardroom contains four things:

  • The inventory position, how many systems, how many high-risk, how many still awaiting classification. 
  • The obligations calendar, which duties bite when, against the post-Omnibus dates. 
  • The exceptions report, systems operating outside policy, residual risks accepted above threshold, incidents. 
  • Decisions sought, stated as decisions, not buried in narrative.

Quarterly is a defensible rhythm for most portfolios. But frequency matters less than one simple thing. Minutes must exist. A board that was demonstrably informed is itself an accountability artefact, and one that regulators, and increasingly insurers, ask to see.

The policy hierarchy

Documentation in mature organisations comes in layers, and the AI file should respect the same discipline.

At the top, one short AI policy. The organisation’s position on what it will and will not deploy. Who owns compliance. And the standing principle that no system enters use unregistered and unclassified. Two pages. It should change rarely.

Beneath it, the standards. The classification procedure from Chapter 5. The risk management procedure from Chapter 6. The data governance procedure from Chapter 7. The procurement standard this chapter reaches shortly. And beneath those, the system-level records the standards generate.

The hierarchy earns its keep in both directions. Downward, a new system’s compliance path is prescribed before anyone gets to argue about it. Upward, an auditor can trace any system record to the standard that required it, and the standard to the policy that required the standard. 

That trace is what a functioning management system looks like from the outside. Organisations already running ISO/IEC 42001 will recognise this as their existing document architecture with the Act’s obligations mapped in. Extend it. Do not build a duplicate next door.

The AI inventory as the master record

If the file has a spine, it is the inventory. Every other artefact in this book attaches to a system, and the inventory is where the systems are counted. Which makes it the record that proves the other records are complete. 

A risk file covering nine systems is excellent evidence about nine systems, and perfectly silent about the tenth one nobody registered. Auditors know this. It is why serious inspections begin with inventory completeness, not with the glamorous documents.

A usable inventory row carries the following:

  • System identity and version. 
  • Business owner and system-level compliance owner.
  • The role analysis, provider, deployer or both, remembering from Chapter 1 that roles attach per system and can flip with conduct. 
  • The classification outcome, dated, linked to its Chapter 5 record. Deployment contexts and affected populations. 
  • Upstream dependencies, the vendor, the underlying model, the data sources. 
  • Lifecycle status. 
  • Obligations triggered, with their dates.

The inventory is also where shadow AI comes to die. One standing rule does the work: procurement, IT onboarding and expense approval all route AI acquisitions through registration. 

That rule is worth more than any detection tooling, because it converts the unregistered system from an unknown into a policy breach with a named breacher. People behave differently when the breach has their name on it.

Keep the inventory versioned. A point-in-time export proves what the organisation knew and when. That proof matters precisely in the scenarios you hope never to need it for.

The deployer’s Article 26 duties

Most organisations reading this book are deployers for most of their portfolio, and Article 26 is their operative chapter of the Act for high-risk systems. The duties are concrete, and each one generates evidence the inventory should point to [1, Art 26].

Deployers must take appropriate technical and organisational measures to use the system in accordance with its instructions for use. The instructions are the provider’s Article 13 deliverable. 

The deployer’s first artefact is the record that somebody read them and translated them into operating procedure. And deviation from the instructions is not merely risky. It is one of the conducts that can convert a deployer into a provider, with the full Chapter III burden arriving in the conversion. An expensive way to learn to read manuals!

Deployers must assign human oversight to natural persons with the necessary competence, training and authority. Three nouns, three evidence items:

  • The assignment record. 
  • The training record. 
  • The authority, which means the overseer can genuinely intervene, or decline the system’s output, without career consequences. 

An overseer who cannot override is a spectator, and the file will show it.

Where the deployer controls the input data, it must ensure that data is relevant and sufficiently representative for the intended purpose. Chapter 7’s discipline arriving at the deployer’s desk, in miniature, scoped to the inputs the deployer actually controls.

Deployers must monitor operation against the instructions for use. Inform the provider or distributor when they have reason to think a risk within the meaning of the Act has materialised. Suspend use and inform the market surveillance authority in the serious cases the Article specifies. 

And keep the automatically generated logs under their control for a period appropriate to the purpose, at least six months, unless other law provides otherwise.

That log retention line is the quiet compliance failure in waiting, so it gets its own paragraph. Vendor defaults frequently purge logs on much shorter cycles. Ninety days is common. 

The duty sits with the deployer, not the vendor, and the vendor’s default settings are not a defence. Check the retention setting on every high-risk system before its obligations bite. Then record that you checked.

Two duties reach beyond the compliance team. Before putting a high-risk system into service at a workplace, deployers who are employers must tell workers’ representatives and the affected workers that they will be subject to the system. And deployers whose high-risk systems make or assist decisions about people must, where Article 26(11) applies, inform those people [1, Art 26(11)]. 

Both duties have GDPR cousins, and both are best discharged through the transparency machinery your privacy programme already runs. Extended, not duplicated.

Article 26’s duties follow the deferred timeline, 2 December 2027 for Annex III systems [4]. The deferral is the deployer’s preparation window, and the preparation is unglamorous. Read the instructions. Appoint the overseers. Fix the log retention. Build the worker notification. None of it improves by waiting.

Fundamental rights impact assessments, for those who need them

Article 27 survived the Omnibus, despite industry lobbying that it merely duplicates the GDPR’s impact assessment machinery [16]. The concession went the other way: the final text expressly lets the FRIA cross-reference or incorporate the DPIA, with an AI Office questionnaire template to make the join practical. 

It applies to a defined subset of deployers, before first use of a high-risk system:

  • Bodies governed by public law. 
  • Private entities providing public services. 
  • Deployers of the credit-scoring and life and health insurance pricing systems in Annex III points 5(b) and (c) [1, Art 27(1)].

If you sit outside those categories, the FRIA is not your duty. Nothing stops you borrowing its structure for Chapter 6’s fundamental rights analysis, though, and mature programmes do exactly that.

For those inside the categories, the assessment must describe the deployer’s processes in which the system will be used:

  • The period and frequency of intended use. 
  • The categories of people and groups likely to be affected. 
  • The specific risks of harm to those categories. 
  • The human oversight measures. 
  • What will be done if the risks materialise [1, Art 27(1)(a)–(f)].

Once completed, it is notified to the market surveillance authority. Where the obligations are already met through a GDPR data protection impact assessment, the FRIA complements it. In practice, that means one integrated assessment with the fundamental rights sections extended beyond data protection. 

The efficient implementation is a single template owned jointly by the DPO and the AI compliance owner. The inefficient one is two teams assessing the same deployment in parallel vocabularies, and then discovering, in front of an examiner, that their two documents disagree.

What the auditor asks for

The evidence set for this chapter is organisational rather than technical. Which makes it the easiest to fake badly and the hardest to fake well.

  • The AI policy, dated, board-approved. 
  • The appointment records, the accountable owner, the working group’s terms of reference, the system-level owners. 
  • The inventory, versioned, with the completeness controls routing acquisitions into it. 
  • The board packs and minutes showing the reporting rhythm actually ran. The literacy programme records under Article 4. 
  • Per high-risk system deployed: the instructions-for-use review, the oversight assignments with training and authority evidence, the log retention confirmation, the monitoring records, the worker and affected-person notifications, and the FRIA with its authority notification wherever Article 27 applies.

Ownership is the theme running through all of it. Every artefact in this list carries a name. An auditor reading the file should never once have to wonder whose job something was.

A specimen inventory row and FRIA outline appear in the appendices. The working versions, with the full inventory schema, board pack template and integrated FRIA-DPIA instrument, are part of the template collection described at the back of this book.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 9: AI Literacy, Article 4

Article 4 is the AI Act’s most widely applicable provision and its most casually dismissed one. It applies to every provider and every deployer regardless of risk tier, it has been in application since 2 February 2025, which is longer than almost anything else in the Act, and it gets routinely waved away as the soft obligation that nobody enforces. The dismissal misreads both how the provision works and what it does inside a compliance file.

This chapter covers what the obligation requires, what the Digital Omnibus did to it, what regulators have signalled along the way, and how to build and evidence a literacy programme proportionate to your organisation. 

The evidencing sections double as the manual for the AI Literacy Compliance Kit, which packages the templates this chapter describes.

The obligation as enacted

As adopted, Article 4 requires providers and deployers of AI systems to take measures to ensure, to their best extent, a sufficient level of AI literacy among their staff and other persons dealing with the operation and use of AI systems on their behalf [1, Art 4]. 

The measure is calibrated to circumstances: their technical knowledge, experience, education and training, the context the systems are used in, and the persons or groups on whom the systems are used.

Four features of that drafting deserve attention. 

The addressees come first: providers and deployers alike, with no risk-tier threshold anywhere, so a company whose only AI is one procured chatbot is addressed. 

Then the population, which covers staff and anyone else operating AI on the organisation’s behalf, a phrase that reaches contractors and outsourced functions. 

Then the standard, sufficient literacy to the organisation’s best extent, which is an effort-and-proportionality formula rather than a fixed competence bar. 

And finally the calibration factors, which do real work in practice: they make a single generic e-learning module for everyone a literal misreading of the provision, because the provision itself orders differentiation by role, by context and by affected persons.

AI literacy has its own definition in Article 3(56): the skills, knowledge and understanding that allow informed deployment of AI systems, together with awareness of the opportunities, risks and possible harms of AI [1, Art 3(56)]. 

The definition is functional rather than academic, tied to what the person actually does with the system.

One structural point explains the quiet enforcement history. Article 99’s fine tiers do not list Article 4 at all. Enforcement runs indirectly, through national penalty regimes, through market surveillance powers, and most concretely through the provision’s evidential role when something else goes wrong. That last route is the one this chapter keeps returning to.

What the Omnibus did to Article 4

The provision’s recent history is a three-stage fight, and knowing the stages matters because commentary from each one still circulates as though it were current.

Stage one was the Commission’s November 2025 proposal, which removed the operator obligation outright [17] and shifted responsibility to the Commission and the Member States, who would promote literacy and encourage proportionate training from an institutional distance. The stated rationale was burden reduction for smaller enterprises, plus avoiding overlap with the competence requirements that attach to high-risk systems anyway [18].

Stage two came from the Parliament, which pushed back hard. The rapporteurs’ amendments reinstated an operator obligation with the verb changed from ensure to promote [19], converting an obligation of result into an obligation of conduct, and added an institutional support duty on the Commission and Member States alongside it. The plenary text adopted in March 2026 required providers to take steps to help improve AI literacy among staff and others operating systems on their behalf [20]. The data protection authorities opposed the deletion throughout [21].

Stage three settled it. The final text, in force since 27 July 2026, keeps the obligation on operators and fixes its verb: providers and deployers shall take measures to support the development of AI literacy of their staff and others operating AI on their behalf, with the original calibration factors retained word for word [1, Art 4]. 

Then it adds a sentence the Commission’s deletion attempt paid for: the obligation does not require providers or deployers to guarantee any specific level of AI literacy of any individual. 

The Parliament’s line won. The obligation survives as a duty of conduct, not of result, and the institutional layer arrived alongside it, the Commission and Member States must support operators’ efforts, SMEs in particular, with practical examples published on the Commission’s single information platform, and the AI Board adopting recommendations with common objectives [1, Art 4(2), (3)].

Three things came through the fight unchanged. The period from February 2025 to mid-2026 ran under the ensure formulation, so the record of what your organisation did in that window gets judged against the stricter wording. 

The competence requirements attached to high-risk systems survive completely untouched: deployers must assign human oversight to persons with the necessary competence, training and authority, and support them, under Article 26(2), while providers must design systems so oversight persons can understand, monitor and intervene, under Article 14 [1, Arts 26(2), 14]. 

In any case, the definition of literacy in Article 3(56) stands, so the concept keeps its content even where the obligation’s edge softened.

What regulators have signalled

Formal enforcement of Article 4 has not materialised, which is consistent with its penalty structure. The signalling, however, has been consistent on four points.

Proportionality is real rather than rhetorical. The Commission’s published answers on Article 4 emphasise measures appropriate to context and role. There is no certification requirement, no mandated curriculum, and no fixed number of training hours anywhere in the picture.

Documentation is the actual deliverable. The same guidance makes clear that organisations should be able to demonstrate the measures they took. An unrecorded training conversation satisfies nobody, because the obligation’s practical output is a record.

Generic tooling fails as risk rises. The signal running through the Commission material and the Article 26 structure alike is that literacy converges with competence for consequential systems. The person overseeing a credit scoring model needs training on that specific model, its limits and its failure modes. A general awareness deck does not get there.

And the softening is itself a signal worth reading. The Omnibus fight shows the Commission regards blanket operator-level literacy mandates as disproportionate for low-risk contexts, while the Parliament and the data protection authorities regard workforce literacy as load-bearing. 

Where systems are consequential, enforcement attention will follow the second view despite the softened verb, because oversight competence under Articles 14 and 26 is where literacy failures become actionable.

Why literacy still pays, whatever the final text

What follows is the honest business case, stripped of the compliance-theatre version.

The competence obligations are unconditional. If you deploy any high-risk system, now or by the December 2027 date, Article 26(2) requires competent, trained and supported oversight personnel, and your literacy programme is where that competence gets built and evidenced. 

A programme started only when the high-risk obligations bite produces oversight staff with a few weeks of familiarity behind them. Chapter 3’s duration argument applies here with full force.

Literacy records are exculpatory evidence. When an AI-involved incident happens, a discriminatory screening outcome, a chatbot commitment upheld in court, a data leak through a generative tool, the first internal question is who knew what the system could do, and the first external question is whether the organisation equipped its people. 

A dated training record on the specific system, covering its limits and its escalation routes, marks the difference between an individual error and an organisational failure. This is ordinary negligence and vicarious liability logic, and it operates regardless of Article 4’s softened wording.

Counterparties are already asking. Enterprise procurement questionnaires and insurers’ proposal forms increasingly include workforce AI training among their questions, and “the obligation was softened” is not an answer any procurement scoring matrix rewards.

Misuse is the cheapest risk you can reduce. Most organisational AI harm in practice is not exotic model failure. It is staff pasting confidential material into public tools, over-trusting fluent output, or failing to recognise a decision they were supposed to check. 

Training against these patterns is among the highest-return controls in the entire programme, which is why it would belong in this book with or without an article number attached to it.

Proportionality by role

The calibration factors in Article 4 and the competence duties in Article 26 point in the same direction: literacy is a matrix, not a module. Six populations recur, each with the content weighting that fits it.

Everyone

This is the floor tier: what AI systems the organisation runs and permits, the acceptable use policy, the confidentiality rules for external tools, the failure modes of generative systems including fabrication, and where to escalate. Keep it short and concrete, and renew it when policy changes.

Operators of specific systems

These are people whose daily work runs through a particular system, recruiters inside the screening tool, underwriters with the pricing model, agents working beside the chatbot. The content is system-specific: the intended purpose, the instructions for use translated into working language, the known limitations, the outputs that require checking, the escalation triggers. 

This tier gets built per system from the provider’s documentation, which is one more reason Chapter 12’s instructions-for-use discipline matters.

Oversight personnel

This is the Article 26(2) population, and they get everything in the operator tier plus the oversight design itself: what they are empowered to override, how automation bias operates on them personally, the intervention and stop procedures, and the record their interventions should leave behind. Depth here is a legal requirement for high-risk systems rather than a training preference.

Builders and integrators

Developers, data scientists and engineers configuring or building systems need the Act’s architecture as it lands on design work: classification consequences, the Chapter III requirements read as engineering constraints, the logging and documentation duties, and the Article 25 perimeter from Chapter 2, because they are the population most able to walk the organisation across that line without anyone noticing.

Procurement and vendor management

This is the population that signs AI into the building. Their content covers the definitional screen from Chapter 2, the prohibition red flags from Chapter 4, the documentation to demand from vendors, and the contract points Chapters 11, 14 and 15 collect. One trained procurement lead prevents more compliance debt than most controls in this book.

Leadership

Enough to govern, and no more: the risk-tier logic, where accountability sits, what the organisation has deployed and why, the penalty and liability exposure, and what the compliance calendar demands of the budget. An hour honestly spent, once a year, covers it.

Two calibration factors get missed most often. The first is persons on whom systems are used: a workforce deploying AI on vulnerable customers needs training tuned to that population’s risks, per the express words of Article 4. 

The second is other persons operating on your behalf: the outsourced call centre running your chatbot sits inside your literacy perimeter, which in practice means a contractual training obligation with evidence flowing back to you. The Kit provides that clause.

Building the programme

The build sequence is short, and it runs off artefacts you already hold if you have followed this book in order.

Start from the inventory. The Chapter 8 system inventory, filtered by classification, tells you which systems need system-specific training and which populations touch them. A literacy programme built without the inventory trains the wrong people on generalities, and that is the single commonest failure in this whole area.

Map roles to tiers. Produce a one-page matrix: population, systems touched, tier, content owner, cadence. This document is the programme’s spine and the first thing to show anyone who questions whether your measures are proportionate, because it is the proportionality analysis itself, visibly performed.

Source or build content per tier. The floor tier can be built in-house or bought off the shelf. System tiers get assembled from provider documentation and internal procedure. Oversight tiers are built with the system owner in the room. Buying one generic course for all six populations reproduces exactly the one-size-fits-all reading that the provision’s own words reject.

Deliver, and capture completion at delivery. Whatever the channel, live, recorded or written, the completion record gets generated in the same motion as the training itself, or it will never be generated at all. Nobody backfills attendance records well.

Refresh on triggers, not anniversaries. A new system, a changed system, a new population, an incident, a policy change, a regulatory change: each one fires a refresh. An annual all-hands refresher makes a reasonable floor, but the trigger list is what keeps the programme responsive instead of ritual.

For a fifty-person deployer with no high-risk systems, the honest total comes to a floor module, two or three system-specific briefings, a procurement briefing, and a matrix that fits on one page. Days of work, not months. The programme scales with the inventory, and that scaling is precisely the proportionality the provision asks for.

Measuring and evidencing it

The programme’s output is a small, stable evidence set. Five artefacts cover it, and the Kit provides each one as a template.

  • The literacy policy runs to a page or two, stating the organisation’s approach, the tier structure, the ownership and the cadence. 
  • The role-to-tier matrix is the proportionality analysis already described. 
  • The training log records each person, each module, each date, with the module version included, because content changes over time and the log must show who was trained on what. 
  • Attestations convert attendance into a record with a name on it, a signed confirmation of completion and understanding, and this is where a book, this one included, can sit inside a programme, with a reading attestation recording the chapters covered and the date. 
  • Assessment and effectiveness records: a short quiz suffices for the floor tier, while oversight tiers need evidence the person can actually run the intervention procedure. Best of all are records showing the training changed behaviour, escalations made, policy queries raised, near-misses reported. 

Effectiveness evidence is the layer most organisations never reach and the one that most impresses an examiner, because it shows the programme is a control rather than a ceremony.

Measurement follows the same logic in three layers. Coverage means completion rates against the matrix. Currency means the proportion of records still within their refresh window. Effect means the behavioural indicators above. Report those three numbers to whoever owns AI governance under Chapter 8, on the same cadence as the rest of the governance pack.

What an auditor will ask for

For this chapter’s material, expect requests for the following:

  • The literacy policy and the role-to-tier matrix. 
  • The training log, reconciled against the current staff list and against the inventory, meaning every oversight person for every high-risk system has been trained on that system and the record is in date. 
  • The content itself, versioned, for spot review. 
  • Attestations for whatever population gets sampled. For outsourced operators, the contractual training clause together with the evidence that came back. 
  • For the pre-Omnibus period, records showing measures were taken while the ensure formulation applied. 
  • For any incident in the file, the training records of the people involved, because that is the first place an examiner, a claimant or your own insurer will look.

Literacy equips the people, and the chapters before this one built the structures they work inside. Part III now turns from governance to the technical safeguards, beginning with the documentation package of Article 11. 

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Part III — Technical Safeguards

Part II built the governance layer: the decisions about what you run, how risky it is, who answers for it, and what your people actually know. Part III descends into the machinery.

Articles 11 to 15 of the Act, with Article 50 alongside them, regulate the system itself. What gets written down about it. What it records as it runs. What it tells the people using it. How humans stay in command of it. And how well it must perform.

The chapters ahead share one premise, worth stating once. The technical requirements read like engineering obligations, and engineers must indeed implement them. But every single one of them is discharged through documents and records, and the documents are what an authority examines. A system can be superbly engineered and completely indefensible, because nothing evidences the engineering. It happens more often than engineers like to hear.

Part III is therefore written for the person who must make the evidence exist. What each artefact is, who builds it, what it must contain, and how it stays alive after the launch that produced it.

One scoping note applies throughout. The requirements in Articles 8 to 15 attach to high-risk systems, with application dates of December 2027 for Annex III systems and August 2028 for Annex I products under the amended calendar. 

The transparency duties of Article 50, covered in Chapter 14, reach much further and arrive much sooner, from August 2026. Readers with no high-risk systems should still walk through Part III rather than skip it. 

The artefacts it describes are exactly what enterprise customers now demand contractually from every AI vendor, whatever the risk tier. Your regulator may leave you alone for years. Your biggest customer’s procurement team will not.

Chapter 10: Technical Documentation, Article 11 and Annex IV

Ask a compliance officer what Article 11 requires and the answer arrives instantly: technical documentation for high-risk systems. Ask what the documentation is for, and the answers go quiet.

The purpose is precise, and it dictates everything about how the documentation should be built. Article 11 documentation exists to demonstrate, to a competent authority or notified body holding it in their hands, that the system complies with the requirements of Chapter III Section 2 [1, Art 11(1)]. 

It is a proof, written for a sceptical stranger, assembled before the system reaches the market, and maintained for as long as the system lives.

That purpose settles three arguments before anyone starts them. The documentation is not the engineering wiki, though it draws on it. A wiki explains the system to insiders. Annex IV documentation demonstrates conformity to outsiders, and the two audiences forgive entirely different sins. 

It is not marketing collateral either, and here accuracy carries legal weight, because supplying incorrect or incomplete information to authorities has its own fine tier. The Act, in other words, has a specific price list for optimism. 

And it is not homework you can defer. The provider must draw it up before placing the system on the market, and keep it at the authorities’ disposal for ten years afterwards [1, Arts 11, 18]. 

Ten years is longer than most of the systems will live, longer than most of the teams will stay, and considerably longer than anyone’s memory of why a design decision was made. Which is rather the point of writing things down.

Who owes technical documentation

The duty sits on the provider of a high-risk system. Deployers do not owe Annex IV documentation, but they should read this chapter anyway, for two reasons. The instructions for use they receive under Article 13 are generated from this package. 

And deployer due diligence increasingly means knowing what a provider’s documentation should contain, well enough to notice when it does not. Importers and distributors, meanwhile, must verify the documentation exists before releasing a system, per their verification roles from Chapter 1.

Two proportionality mechanisms soften the duty at the edges. SMEs, including start-ups, may provide the Annex IV elements in a simplified form, using a template the Commission is to establish [1, Art 11(1)]. The Omnibus extended this proportionality approach to small mid-cap enterprises as well [4]. 

And where a high-risk system belongs to a product already documented under Annex I legislation, the Act contemplates one set of documentation covering both regimes [1, Art 11(2)]. For regulated-product manufacturers, that is an invitation to extend the existing technical file rather than build a parallel one, and invitations like that are rare enough to accept.

A warning about the simplified form, though. Simplified form is not simplified substance. The SME template reorganises what must be shown. The requirements of Section 2 still apply in full, and the demonstration must still succeed. 

Treat the simplification as formatting relief, nothing more. And be cautious about building to a simplified form the Commission has not yet published, because building to a template that does not exist is a hobby, not a strategy.

The package, section by section

Annex IV specifies the contents in nine parts [1, Annex IV]. Here is each part in operational terms: what it is, who holds the source material, and where the effort concentrates.

General description of the system. Intended purpose, provider identity, system version and its relationship to previous versions, how the system interacts with hardware and other software including other AI systems, the forms in which it reaches the market, the hardware it runs on, photographs or descriptions where the system is embedded in a product, the basic user interface, and the instructions for use for the deployer. 

The source is product management. The effort note: the intended purpose stated here must match the classification record, the marketing, and the instructions for use, word for word. Divergence between these documents is among the first inconsistencies an examiner finds, partly because it is the easiest thing to check and partly because it is so reliably there.

  • Detailed description of elements and development process. 
  • The methods and steps of development, including where third-party tools or pre-trained models were used. 
  • The design specifications, the system’s general logic, the algorithms, the key design choices with their rationale and assumptions, including assumptions about the people the system will be used on. 
  • The architecture, and how the components feed each other. 
  • The data requirements, describing training methodologies and datasets, their provenance, scope and characteristics, the labelling and cleaning procedures, in effect the datasheets Chapter 7 built. 
  • The human oversight measures assessed under Article 14. Any pre-determined changes. The validation and testing procedures, the metrics for accuracy and robustness, the test logs and reports, dated and signed. 

The source is engineering and data science, translated into language an outsider can follow. 

The effort note: this is the heaviest part of the package, and it cannot be written retrospectively with integrity, because it documents a process. A process either left records as it ran or it did not, and no amount of eloquence in 2027 fixes silence from 2025.

Monitoring, functioning and control. The system’s capabilities and limitations, the expected accuracy for the intended purpose, foreseeable unintended outcomes and sources of risk, the oversight measures as actually deployed, and the specifications for input data. Source: engineering plus the risk file. 

This part is where honesty about limitations lives. A documentation package with no stated limitations is a red flag, not an achievement. Every system has limitations. A file that lists none has simply not looked.

Appropriateness of the performance metrics. Why the metrics chosen in Part Two are the right metrics for this purpose and this population. It is short, it is reasoned, it is frequently skipped, and it is the first question a technically literate examiner asks. Draw your own conclusions about that combination.

The risk management system

A description of the Article 9 process as applied to this product. In practice, a pointer into the living risk file, with a summary of its current state. Chapter 6 built it. Here it gets exhibited.

Lifecycle changes

The running changelog of relevant changes made through the system’s life, fed by Part Two’s pre-determined-changes description and the version discipline below.

Standards applied

The harmonised standards applied in full or in part, and where none were applied, a description of the solutions adopted to meet the Section 2 requirements instead. Until the harmonised standards actually land, most packages will carry the second limb, which makes the descriptions longer, not shorter. The absence of a standard was never going to mean less writing.

The EU declaration of conformity

A copy of the Article 47 declaration, once conformity assessment concludes. Chapter 16’s subject.

Post-market monitoring

A detailed description of the Article 72 monitoring system and its plan. Chapter 18’s subject, present in the package from day one, because the plan for watching the system must exist before there is a market to watch it in.

Read as a whole, the Annex demands the biography of a system. What it is, how it came to be, what it was fed, how it was tested, what it cannot do, and how it will be watched. 

Assembling that biography is a coordination problem across product, engineering, data, risk and legal, which is why the ownership section below matters more than any template ever will.

Living documentation versus the folder that dies

Article 11 contains four words that defeat the standard corporate approach: kept up to date.

The standard approach produces a binder. A heroic one-off assembly effort in the weeks before a deadline. Handsome, complete, celebrated at a team lunch, and dead on arrival, because the system it describes changes the following sprint and the binder does not. 

Somewhere in your organisation there is probably such a binder already, describing a system as it stood on one proud afternoon in the past. It is not documentation. It is a portrait!

The failure is structural. Documentation assembled as a project has no mechanism for hearing about change. Documentation maintained as a process does. The difference between the two is not the template. There are five design choices that lead to such results.

Single source, generated exhibits

The authoritative content lives in one maintained repository, and the submission-ready package is generated from it rather than edited by hand. Whether that repository is a document management system, a structured wiki or a compliance platform matters far less than the one rule: nothing is written twice. 

Anything written twice will eventually disagree with itself, usually at the least convenient moment available.

Sections owned by the people who own the facts

Part Two’s data description belongs to whoever owns the training pipeline. Part Three’s limitations belong to whoever owns evaluation. Part One belongs to the product. Compliance owns the whole, the consistency between sections, and the calendar. But a package written entirely by compliance describes a system that compliance did not build, and it reads that way to anyone who has read more than one.

Version control on the documentation itself

Each release of the package is identified, dated, and mapped to the system version it describes, with superseded versions retained. The ten-year duty covers the system’s history, not just its final state, and an authority investigating an incident asks for the documentation as it stood at the time of the incident. Not the improved edition prepared afterwards.

Update triggers, named and watched

The documentation changes when the system changes. Model retraining or replacement, dataset changes, new purposes or populations, interface changes touching oversight, changed input specifications, incidents and their corrective actions, new standards applied, pre-determined changes taking effect. 

Each trigger needs a listener, and the cheapest listener is a gate in the release process: no deployment of a triggering change until the affected sections are updated, reviewed and re-versioned. 

Where release engineering is disciplined, this single control keeps the whole package alive at near-zero marginal cost. The engineers will grumble for a month and then forget the gate exists, which is exactly what success looks like.

Divergence review

On a quarterly or release cadence, a short reconciliation. Does the package still describe the system in production? Do the intended purpose, the classification record, the instructions for use and the marketing still say the same thing? Are the test reports still the tests of the current model? 

Twenty minutes with a checklist. It is the difference between discovering drift yourself and having it discovered for you, and only one of those conversations happens on your terms.

One further cost of the binder approach, worth naming for the budget conversation. Reconstruction is more expensive than maintenance, and visibly less credible. Chapter 3’s timestamp argument applies here with particular force, because Part Two documents a development process, and a development narrative written eighteen months after development, in one authorial voice, on one date, announces itself. It reads like a witness statement drafted by someone who was not there. Examiners have read many of those, and they recognise the genre.

The operating model on one page

Here is the whole system, stated plainly enough to implement on Monday.

Keep one repository. All the authoritative Annex IV content lives there, organised into the nine parts. Nowhere else. If the same fact exists in two places, one of them is wrong, or will be soon.

Give each part an owner. The owner comes from the team that produces the facts. The training pipeline team owns the data description. The evaluation team owns the limitations. 

Product owns the general description. Compliance does not write these sections. Compliance owns three things: assembling the package, checking the sections agree with each other, and keeping the calendar.

Version everything. Each release of the package gets an identifier, a date, and a statement of which system version it describes. When a new version replaces an old one, keep the old one. 

Article 18’s ten-year duty covers the system’s history, and an authority investigating an incident wants the documentation as it stood on the day of the incident, not the polished edition you produced afterwards.

Put a gate in the release process. The trigger list from the previous section becomes a rule: a change that affects the documentation does not ship until the affected sections are updated. 

Real life will occasionally demand an exception, a release that truly cannot wait. Fine. Then a named person accepts the gap in writing, with a date by which it will be closed. An exception with a name and a deadline is a managed risk. An exception without them is just the gate quietly dying.

Run a short reconciliation on a fixed rhythm. Quarterly, or at each release. Does the package still describe what is running in production? One line in the minutes to say the check happened. Twenty minutes, four times a year.

And generate the instructions for use from the same repository as the package. Never maintain them separately. The two documents must always agree, and documents kept apart drift apart. It is practically a law of nature.

One more thing, for providers who have not yet reached the December 2027 line. Not all nine parts can be written now, and they should not be. The parts split into three groups by when their content becomes available.

Parts One, Two and Three describe the system and its development. That content is being created right now, while you build, and it is cheapest to capture in the moment. Parts Seven and Eight depend on harmonised standards and conformity assessment, which have not happened yet, so they sit later on the calendar. Parts Five, Six and Nine describe processes, risk management, lifecycle changes, post-market monitoring, and processes need time to run before there is anything to describe.

So the plan writes itself. Open the repository now. Fill it with what development is generating anyway. Let the process parts accumulate as the processes run. Then the 2027 assembly becomes an editing job instead of an excavation. 

Archaeology is fascinating work. But… You do not want to be doing it on your own product, under a deadline. 

What an auditor will ask for

For this chapter’s material, expect the following requests: 

  • The current technical documentation package for a sampled high-risk system, and then the version that was current at a specified past date, which is the request that finds out whether your version control is real. 
  • The mapping between package versions and system versions.
  • The section ownership record, plus evidence the owners actually maintain their sections, edit histories serve nicely. 
  • The trigger list and the release gate in operation, sampled against the changelog: the examiner picks a material change from Part Six and asks for the documentation update it triggered, with dates. 
  • The reconciliation between the intended purpose in Part One, the classification record, the instructions for use and the public marketing. 
  • For SMEs and small mid-caps on the simplified form, the form itself once published, completed, together with the underlying evidence it summarises. 

The simplification changed the paperwork. It never removed the duty to hold the proof. The next chapter turns to what the system must record about itself as it runs: the logging obligations of Article 12.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 11: Record-Keeping and Logging, Article 12

The documentation package from Chapter 10 describes a system before it meets the world. Logging is the other half of the evidentiary pair, the record the system generates about itself while it meets the world. Article 12 requires high-risk systems to allow for the automatic recording of events, logs, over the system’s lifetime [1, Art 12(1)].

The division of labour between the two is worth holding onto. Documentation proves the system was built compliantly. Logs prove what it actually did on a particular Tuesday. And when an incident, a complaint or an inspection arrives, the Tuesday is what everyone wants.

Article 12 is also where the lawyer-engineer translation problem reaches its sharpest point in the whole Act. The legal text specifies purposes. Engineers need fields, formats and retention rules. 

Most organisations resolve that gap badly in one of two directions. Either legal writes a policy nobody can implement, forty pages of aspiration that never touches a codebase. Or engineering logs whatever the platform logs by default and calls it compliance, which is roughly the same as calling whatever is in the fridge dinner. 

This chapter runs the translation properly, and it ends with the artefact that carries it: a one-page logging specification per system, written by both disciplines, implementable by one and defensible by the other.

What the Act requires, precisely

Three layers of obligation stack up, and they address different actors.

The design duty sits on the provider. High-risk systems must technically allow automatic recording of events over their lifetime. This is a product requirement, meaning a high-risk system that cannot log is non-conforming as designed, in exactly the same way as one that cannot be overseen. The logging capability gets assessed in conformity assessment and described in the technical documentation, in Part Two of the Annex IV package.

The purpose specification shapes what the logging must capture. The recording must ensure a level of traceability appropriate to the system’s intended purpose [1, Art 12(2)], and in particular it must enable three things. Identifying situations that may result in the system presenting a risk, or in a substantial modification. Facilitating post-market monitoring under Article 72. And supporting the deployer’s monitoring of the system’s operation under Article 26(5).

Notice what this drafting does. It defines sufficiency by function, not by field list. Logging is adequate when those three consumers can do their jobs from it. That functional test is the one this chapter operationalises below.

There is one place where the Act does write fields, and that is biometrics. For remote biometric identification systems, the logs must record, at minimum, the period of each use with start and end date and time, the reference database the input was checked against, the input data that produced a match, and the identity of the natural persons who verified the results under Article 14(5) [1, Art 12(3)]. 

Even organisations nowhere near biometrics should read that list once, because it reveals the legislator’s mental model of a good log. Who ran what, against what, with what result, and which human checked it. That model travels well beyond biometrics.

Then come the custody duties, and they split by role. Providers must keep the logs their high-risk systems generate, to the extent those logs are under their control, for a period appropriate to the intended purpose and at least six months, unless other law provides otherwise [1, Art 19]. 

Deployers carry a mirrored duty for the logs under their control, again with the six-month floor [1, Art 26(6)]. The phrase doing the allocation is “under their control”. In a SaaS deployment, the platform logs sit with the provider, the usage and decision context sits with the deployer, and the contract had better say which is which. 

The specification below forces the question, on the theory that questions are cheaper before signature than after.

What to capture: the functional method

Because the Act defines logging by what it must enable, the design method is to walk through the consumers, the three statutory ones plus two practical ones, and derive the fields from their questions.

The risk spotter

Article 12(2)(a) wants the logs to reveal emerging risk and creeping modification. The fields that follow: 

  • Input characteristics sufficient to detect distribution drift. 
  • Output distributions over time. 
  • Confidence or score values, not just the final classifications. 
  • Error and exception events. 
  • The model version on every transaction, because without it drift is indistinguishable from a deployment. 
  • Threshold and configuration values in force at the time of each decision, because a system silently retuned is Chapter 2’s substantial modification problem wearing an operational disguise.

The post-market monitor

Article 72 monitoring, Chapter 18’s subject, needs performance in the field. Volumes, outcome rates, override and complaint correlations, and segment-level patterns wherever the risk assessment identified protected or vulnerable populations. 

The field that follows is enough decision context to aggregate by the segments the risk file cares about. Write that sentence carefully, and write it with data protection counsel in the room, because it invites recording sensitive attributes, and it should usually be satisfied by cohort identifiers rather than raw characteristics.

The deployer’s monitor

Article 26(5) has the deployer watching operation against the instructions for use. The fields that follow: per-transaction records a non-engineer can actually interrogate. The human actions taken on each output, accepted, overridden, escalated, and by whom, which happens to be exactly the evidence Chapter 12’s oversight design needs to prove the oversight is real. Two birds, one log. And timestamps throughout, synchronised to a stated clock.

Then the two consumers the Act implies without naming. The incident reconstructor: when Article 73 reporting or a legal claim arrives, the question is what happened in this specific case, so every logged event needs a stable transaction identifier joining input, system version, output, confidence, human action and downstream effect into one reconstructable chain. 

And the examiner: an authority sampling your logs is testing whether they match the technical documentation’s promises, so the logging described in Annex IV Part Two and the logging actually running in production must be the same logging. One more divergence for Chapter 10’s review to catch.

Two boundaries stop the field list growing forever. 

First, log decisions and events, not everything. Article 12 is about traceability of the system’s operation, not total surveillance of it, and every field you add carries GDPR weight, storage cost and breach surface. A log is also a liability, a fact that becomes vivid at the first data breach. 

Second, build personal data minimisation into the design from day one. Pseudonymise identifiers where reconstruction does not need names, prefer references to payloads, and write the retention rule per field rather than per log.

Retention: the floor, not the number

Six months is a floor with an appropriateness test sitting on top of it, and the honest analysis rarely lands on the floor.

Several factors push longer. 

  • The system’s decision cycle, since a recruitment decision’s consequences unfold over months.
  • The limitation periods for the claims the system could generate, and discrimination and consumer claims run for years. 
  • The observation windows in the post-market monitoring plan. 
  • Sectoral law, with financial services record-keeping as the obvious case.
  • And the ten-year documentation horizon of Article 18, which logs supporting a conformity narrative may need to echo. 

The factors pushing shorter are the familiar ones: storage limitation under data protection law, breach exposure, and cost.

The resolution is a per-system retention schedule with reasons attached. This system, these log categories, these periods, because of these factors, signed and dated. 

Two refinements earn their keep. Tiered retention keeps full transaction detail for the shorter period and aggregated or pseudonymised derivatives for the longer one, which satisfies the monitoring consumers without hoarding raw personal data. 

And a litigation hold trigger overrides the whole schedule the moment an incident or claim crystallises, with a named owner in the specification. Logs deleted on schedule three weeks after an incident is a fact pattern no counsel wants to explain, least of all to a judge who has seen it before.

Access: who reads the logs

Logs are evidence, and evidence has custody rules. Four access questions belong in the design rather than in the aftermath.

Who inside the organisation reads them, and for what

Oversight personnel for their monitoring, engineering for operations, compliance for audit, each with access scoped to purpose. The logs contain personal data, and Article 12’s purposes do not include curiosity.

Who outside can demand them

  • Market surveillance authorities, under Article 21 for providers and through the Article 26 mechanics for deployers. 
  • The provider, from the deployer, where post-market monitoring needs field data, a flow the contract must create because the Act simply assumes it exists. 
  • Litigants, eventually, through disclosure.

What integrity protection they carry

Append-only storage or equivalent tamper evidence, access logging on the logs themselves, yes, logs about the logs, and clock discipline. Logs that could have been edited prove very little, and everyone in the room will know it.

Where they physically sit

In the SaaS pattern this question returns, as most questions do, to the contract. The deployer’s Article 26(6) duty covers logs under its control, so the deployer needs either the logs themselves or a contractual right to timely export. Silence on this point is the single commonest gap I see in AI procurement. The vendor is not hiding anything. Nobody asked.

The one-page logging specification

Now the bridge artefact

One page per high-risk system, co-authored by the system owner and compliance, reviewed by data protection, versioned alongside the technical documentation. Its sections:

System and scope

The system name, the version range covered, whether this is the provider’s or the deployer’s perspective, and the contractual allocation of log custody, in one sentence.

Events captured

The field list, derived by the functional method above, with each field tagged by the consumer it serves, risk, monitoring, deployer, reconstruction. Every field has a reason, and every statutory purpose has fields. Where biometrics are anywhere near scope, the Article 12(3) quartet appears verbatim.

Identifiers and joins

The transaction identifier scheme and what it joins together. The model and configuration version fields. The clock source.

Personal data treatment

Which fields carry personal data, what minimisation was applied, and the pointer to the GDPR record of processing that covers the logging itself.

Retention

The per-category periods with their one-line reasons, the tiering, and the hold trigger with its owner.

Access

The internal roles and their scopes, the external routes, and the integrity measures.

Verification

How the organisation knows the specification is actually true in production. The test performed at deployment, its cadence afterwards, and the last date it ran, with the result.

That last section is what separates a specification from an aspiration. A specification nobody has tested against production is a description of intentions, and the gap between intended logging and actual logging is precisely what incident reconstruction discovers, always at the worst available moment. 

The test itself is cheap. Sample a transaction, pull its chain, check that every specified field arrives populated, file the result. Quarterly, or whenever Chapter 10’s release gate fires on a triggering change.

For lawyers, the specification is what you hand an engineer instead of a policy. For engineers, it is what you hand an examiner instead of a shrug. It is also, quite deliberately, the smallest document in the entire Part III set. The discipline of one page forces the prioritisation that forty-page logging policies exist to avoid.

What an auditor will ask for

For this chapter’s material, expect the following requests:

  • The logging specification for a sampled system, and the technical documentation’s logging description, read against each other. 
  • A sampled transaction reconstructed end to end from the logs, input, version, output, human action, timestamps and all. 
  • The retention schedule with its reasoning, plus evidence that deletion actually runs, because retention promises without deletion evidence are read as neither. 
  • The access model, and the access logs on the logs. 
  • The contractual custody allocation for procured systems. 
  • For deployers, the demonstration that six months of logs under your control exist right now, for every high-risk system in the inventory, which is the single fastest test of whether this chapter was implemented or merely read. 
  • After any incident, the hold record showing the deletion schedule was suspended in time.

Logs record what the system did. The next chapter turns to the humans standing over it: transparency to deployers and human oversight, Articles 13 and 14.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 12: Transparency and Human Oversight, Articles 13 and 14

Articles 13 and 14 form a pair, and the Act’s architecture only works when they are read as one. Article 13 requires the system to be transparent enough for its deployer to use it properly. Article 14 requires it to be overseeable by humans while in use.

The connecting logic is a chain of enablement. The provider designs for oversight and explains the system through the instructions for use. The deployer assigns competent people, Chapter 9’s population, who exercise the oversight the design permits. 

Break any link and the others fail with it, which is why oversight failures usually turn out to be transparency failures or literacy failures wearing a different name.

One point of positioning first, carried forward from Chapter 5. The human in the loop is not a classification defence. A system that scores, ranks or recommends within an Annex III area is high-risk even though a human decides at the end. The human is what Article 14 regulates. The human is not what Article 6(3) exempts. 

So this chapter is not about escaping obligations. It is about discharging the one that attaches to almost every high-risk deployment: making human control real, and provably real.

Article 13: transparency to the deployer

Article 13’s addressee gets misread constantly. This is not public-facing transparency, that is Article 50, and it lives in Chapter 14. Article 13 requires high-risk systems to be designed and developed so that their operation is sufficiently transparent for deployers to interpret the system’s output and use it appropriately [1, Art 13(1)]. 

The measure of sufficiency is functional once again. Transparent enough for the deployer to comply with its own duties.

The operative artefact is the instructions for use. They must: 

  • Accompany the system in an appropriate format and contain concise, complete, correct and clear information, relevant, accessible and comprehensible to deployers [1, Art 13(2), (3)]. 
  • The required contents track the technical documentation in miniature. 
  • The provider’s identity. 
  • The system’s characteristics, capabilities and limitations, including its intended purpose. 
  • The level of accuracy with its metrics, plus the robustness and cybersecurity levels the system was tested against. 
  • Any known or foreseeable circumstance that may lead to risks, the system’s performance for specific persons or groups, and the specifications for input data. 
  • Any pre-determined changes. 
  • The human oversight measures under Article 14, including the technical measures that help interpret outputs. 
  • The computational and hardware resources needed, the expected lifetime, the necessary maintenance. 
  • Where relevant, a description of the mechanisms allowing deployers to collect, store and interpret the logs, which answers Chapter 11’s custody question at source [1, Art 13(3)].

Now treat the instructions as a compliance artefact rather than a manual, because they are doing four jobs at once, and drafting them well means drafting for all four:

  1. They are the provider’s risk transfer. Limitations and foreseeable misuse stated here become the deployer’s problem to manage. Omissions here remain the provider’s problem forever.
  2. They are the deployer’s operating constraint. Article 26(1) obliges deployers to use the system in accordance with them, which means every sentence in the document is a duty somebody inherits. That is also why they must be written for the deployer’s actual reading level, not the engineering team’s.
  3. They are the perimeter of safe customisation from Chapter 2. What the instructions foresee, the configuration ranges, the retraining procedures, defines exactly where Article 25 begins.
  4. They are evidence in every direction. In an incident, the first two documents on the table are the instructions and the logs, and the case usually turns on the distance between them.

For providers, three drafting disciplines follow. Write limitations as operational sentences rather than disclaimers. “Accuracy degrades materially for inputs in languages other than those listed” instructs somebody to do something. “Outputs should be verified” insulates nobody and instructs nobody. 

Generate the instructions from the Chapter 10 repository, so they cannot quietly diverge from the technical documentation. And version them with the system, because instructions describing last year’s model are worse than none at all. They are wrong, and they are being relied upon, which is the worst available combination.

For deployers, the mirror disciplines. Obtain the instructions. Actually read them, which is rarer than it should be. Operationalise them into working procedure and into Chapter 9’s operator training. Then file the mapping between instructions and procedure, because “we used the system in accordance with the instructions” is a claim you will one day need to prove sentence by sentence.

Article 14: the oversight requirement

Article 14 requires high-risk systems to be designed and developed, including with appropriate human-machine interface tools, so that natural persons can effectively oversee them while in use, with the aim of preventing or minimising risks to health, safety and fundamental rights [1, Art 14(1), (2)].

The duty splits across the chain, and the split itself is a design decision. Oversight measures are either built into the system by the provider before market, or identified by the provider as appropriate for the deployer to implement [1, Art 14(3)]. The deployer then assigns the oversight to people with the competence, training, authority and support the task requires, under Article 26.

What the provider must enable is specified as capabilities of the overseeing human, and the list repays exact reading [1, Art 14(4)]. The overseer must be able to properly understand the system’s capacities and limitations, and monitor its operation, including for anomalies and unexpected performance:

  • To remain aware of automation bias, the tendency to over-rely on the system’s output, particularly where the system informs human decisions. 
  • To interpret the output correctly, using the interpretation tools available. 
  • To decide not to use the system at all, or to disregard, override or reverse its output. 
  • To intervene in the operation, or interrupt it through a stop function or similar procedure that brings the system to a halt in a safe state.

One special rule sits on top. For the remote biometric identification systems of Annex III point 1(a), the deployer may take no action or decision on the basis of an identification unless it has been separately verified and confirmed by at least two natural persons with the necessary competence, training and authority, subject to narrow law-enforcement carve-outs [1, Art 14(5)]. 

Two humans, independently, before anything happens to anyone. The legislator’s arithmetic is worth noticing.

The three oversight models

Practice has settled on three architectures, and Article 14 accommodates all of them. The compliance work is choosing deliberately, and matching the choice to the risk file.

Human-in-the-loop

A person acts on every output before it takes effect. The system proposes, the human disposes. The recruiter reviews every ranked shortlist, the underwriter confirms every declined application. 

This is the strongest control at the highest cost, and it has one dominant failure mode: rubber-stamping, where throughput pressure converts review into ritual, and the human adds latency but no judgement. 

In-the-loop is the right model where individual decisions carry high consequence, where volumes permit genuine attention, and where the risk assessment identified error types a competent human can actually catch. 

That last condition is the one nobody checks. A human cannot meaningfully review a fraud score derived from four hundred features, and pretending otherwise is decoration.

Human-on-the-loop

The system acts autonomously, and a person monitors operation, intervening on exception. The human watches dashboards, samples outputs, handles escalations, and holds the stop button. 

This is the right model where volume or speed makes per-decision review impossible. But it is only honest where the monitoring is properly instrumented, drift alerts, anomaly flags, sampled review with quotas, and where the intervention path is fast enough to matter for the harm in question. 

Its failure mode is the empty control room. Monitoring assigned as a duty, resourced as an afterthought.

Human-in-command

A person holds authority over the system’s deployment envelope. When it runs, on which population, within which limits, and whether it runs at all, which is Article 14(4)’s often-forgotten capability. 

This is governance-level oversight, and it is not an alternative to the other two models. It is the layer above them. Every deployment needs someone in command. The only question is whether per-output or exception-based control sits underneath.

Choosing between the models is an output of the Article 9 risk process, not a preference. The method runs like this. 

Take the harm scenarios from the risk file. For each one, ask what a human could realistically detect and reverse, at what point in the flow, within what time. Cost the honest version of that control. Then document the choice, with the rejected alternatives. 

A written record that says “on-the-loop was chosen over in-the-loop because per-decision review at this volume produces rubber-stamping, and sampled review at this percentage with these alert thresholds catches the error classes the risk assessment rates as material” is worth more in an inspection than any amount of interface polish. It shows the oversight was designed against the actual risks, rather than installed as furniture.

Interface-level requirements: no such thing as a simple red button

Article 14 does not stop at policies and training. It requires “appropriate human-machine interface tools”, which means the screen itself is regulated. Each capability the overseer must have translates into something the interface must do.

Understanding and interpretation

The screen should show how confident the system is, not a bare label. A score of 51 and a score of 99 produce the same “decline”, and the overseer needs to see the difference. Where the model class allows it, show what drove the output. And put warnings where the risk arises. If accuracy drops for unsupported languages, the warning belongs on the screen at the moment such an input appears, not on page forty of a manual.

Automation bias

People trust machine output more than they should, so the design must push back. Add friction where the stakes are high: before confirming a decline, the reviewer must open the underlying evidence. 

Avoid anchoring: where feasible, show the case before the score, so the human forms a view first. And give the accept and override buttons equal prominence. That sounds trivial. After an incident, a screen with a large green accept button and a buried override option becomes an exhibit.

Disagreement

Overriding must be easy. A handful of clicks at most, no justification longer than what acceptance requires, and visible confirmation that the override took effect. If disagreeing with the system costs the reviewer ten minutes and agreeing costs one click, you have designed obedience and called it oversight.

Interruption

The stop control must exist, the assigned people must be able to reach it, and pressing it must lead somewhere defined. For a decision system, a safe state means a written manual fallback, who handles the queue and how, not an off switch and a pile of unprocessed cases.

Two evidentiary hooks belong in the interface from day one, because Chapter 11 already built their destination. Every human action on an output, accept, override, escalate, stop, writes to the log with actor and timestamp. 

And the interface version gets recorded alongside the model version, because oversight adequacy is a property of the pair, not of either half alone.

Proving that oversight is real

Fake oversight has a signature, and examiners, claimants and journalists all recognise it. A human formally in the loop, one hundred per cent agreement with the machine, three seconds per case. If your numbers look like that, no design document will save you. Proving the opposite is an evidence problem, and the evidence is cheap if you built the interface described above.

The evidence comes in three layers.

The paper layer

  • The oversight design record for each system: which model you chose, which alternatives you rejected, which risks the choice addresses. 
  • The assignment records: named people, their training on this specific system, and their authority written into role descriptions that match how they actually report. 
  • The mapping against the instructions for use, showing that each oversight measure the provider identified was either implemented or consciously varied.

The behavioural layer

This is where realness lives, because paper describes intentions and behaviour reveals facts. 

  • Track override and rejection rates per overseer. A rate near zero suggests rubber-stamping. Rates that vary wildly between overseers suggest the training failed. 
  • Track time spent per case against a stated floor, and remember that three seconds is not a review. 
  • Track escalations and what happened to them. 
  • Test the stop function, on a date, with the safe state verified, and file the result, this is the fire drill nobody runs. 
  • Sample cases for blind re-review: a second person works the same case without seeing the first decision, then you compare.

The management layer

Report these numbers to whoever governs AI, on the same cadence as everything else. Set thresholds that trigger action. Then make sure at least one action is on the record: a retrained overseer, a redesigned screen, a paused deployment. 

One genuine intervention proves the whole apparatus better than a thousand pages of design, because it shows the loop closes.

For deployers buying systems, this section doubles as a procurement checklist, run in reverse. Does the product show confidence levels, decision drivers and an override button? Does it log what the human did? Can the vendor produce these behavioural metrics? A provider that cannot answer is selling you an Article 26 problem with a licence fee attached.

What an auditor will ask for

For this chapter’s material, expect the following requests: 

  • The instructions for use for a sampled system, checked against the technical documentation for consistency, against Article 13(3) for completeness, and against your operating procedures for implementation. 
  • The oversight design record, rejected alternatives included. 
  • The assignment and training records for the named overseers, reconciled against the current staff list, because people leave and records do not notice on their own. 
  • The interface walkthrough against the Article 14(4) capabilities, with the examiner watching an override and a stop actually performed. 
  • The behavioural metrics for the trailing period, override rates, timing distributions, escalations, plus the governance minutes showing somebody reviewed them. 
  • The stop-test log. 
  • For biometric deployments, the two-person verification records, with the competence evidence behind both names.

Oversight keeps humans in command of the system. The next chapter turns to what the system itself must withstand: accuracy, robustness and cybersecurity under Article 15.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 13: Accuracy, Robustness and Cybersecurity, Article 15

Article 15 is one sentence with three decades of engineering packed inside it. High-risk systems must be designed and developed to achieve an appropriate level of accuracy, robustness and cybersecurity, and to perform consistently in those respects throughout their lifecycle [1, Art 15(1)].

That single sentence hides three separate jobs. Accuracy asks whether the system gets it right. Robustness asks whether it keeps working when conditions turn hostile or strange. Cybersecurity asks whether it resists people actively trying to make it fail. Each job produces its own evidence. 

And one word haunts all three, because for most systems, no harmonised standard yet tells you what “appropriate” means. This chapter covers the three jobs first, then tackles the word.

Accuracy: declare it, then live with the declaration

The Act does not set an accuracy threshold. No article anywhere says a recruitment tool must be 95 per cent correct. Instead, the Act creates a declaration mechanism. The levels of accuracy and the relevant accuracy metrics of a high-risk system must be declared in the instructions for use [1, Art 15(3)].

Understand what this mechanism does, because it is cleverer than it looks. You choose the metric. You measure your own system. You publish the number to your deployers. And from that moment, the number is a promise, with three audiences holding you to it. Deployers build their oversight around it. Market surveillance authorities test against it. And claimants’ lawyers read it aloud in court if the system underperforms it. The legislator did not need to set your threshold. It arranged for you to set it yourself, in writing, in front of witnesses.

So the compliance work sits in three choices, all made before anything gets declared.

Choose metrics that fit the harm, not the marketing. A single accuracy percentage is almost always the wrong declaration. A fraud model that is 99 per cent accurate can still miss most of the fraud, if fraud is rare, which it usually is. 

For any system making decisions about people, the errors have directions, and the directions have different victims. A false positive in fraud screening blocks an innocent customer. A false negative waves a fraudster through. Declare those rates separately. 

The deployer needs both to design oversight, and a blended number is exactly the kind of statement Chapter 10 called material inaccuracy waiting to happen.

Measure on data that resembles deployment. Accuracy on the training distribution and accuracy in the field are two different numbers, and the gap between them is where incidents live. 

Test on held-out data that matches the population, the languages and the conditions your intended purpose describes. Where the risk file flagged specific groups, measure per group. 

Annex IV Part Two already requires you to document performance for specific persons or groups, and a declared average that conceals a known per-group failure will read, in hindsight, as concealment. Because that is what it was.

Say what the number does not cover. The declaration lives in the instructions for use, so Chapter 12’s drafting rule applies here too. Write it as an operational sentence. “Accuracy of X on populations of type Y, measured by metric Z, degrading materially under conditions W” tells a deployer what to do. A bare percentage decorates a sales deck.

One more duty hides in Article 15(1), easy to miss. Performance must be consistent throughout the lifecycle. A declaration that was true at launch and false after six months of drift is a live non-conformity, quietly ticking. This is why Chapter 11’s logging captured output distributions and model versions. 

The accuracy declaration needs a monitoring loop standing behind it, and the post-market plan in Chapter 18 is where that loop reports.

Robustness: the system meets the real world

Robustness is accuracy’s stress test. The Act requires high-risk systems to be as resilient as possible regarding errors, faults or inconsistencies that may occur within the system or its environment, in particular through interaction with natural persons or other systems [1, Art 15(4)].

Translate that into questions an engineer can actually test:

  • What happens when the input is malformed, incomplete, in the wrong language, or from a population the training data barely saw? 
  • What happens when an upstream system sends garbage, or the API everything depends on times out? 
  • What happens when a user does something no designer imagined, which a user will, usually within the first week? 

The honest answers become the test plan. The test results become Annex IV Part Two evidence.

The Act names two acceptable engineering answers: technical redundancy solutions, which may include backup or fail-safe plans [1, Art 15(4)]. In plain terms, the system needs a defined behaviour for its own bad days. 

For a decision system, this connects straight back to Chapter 12’s safe state. When the model fails or degrades, there is a manual fallback, and somebody owns it by name.

Continuously learning systems get a special rule, and it matters far more than its length suggests. Systems that keep learning after deployment must be built to eliminate or reduce, as far as possible, the risk of feedback loops, biased outputs influencing future inputs, which then confirm the bias [1, Art 15(4)]. 

The classic case is a policing or credit model whose own past decisions shape the data it learns from next. The model becomes a machine for proving itself right. If your system learns in production, this paragraph requires you to show how the loop is broken. “We retrain periodically” is not an answer but rather a symptom of a bigger problem.

Cybersecurity: someone will be trying to break it

The third limb assumes an adversary. High-risk systems must be resilient against attempts by unauthorised third parties to alter their use, outputs or performance by exploiting system vulnerabilities [1, Art 15(5)].

Ordinary security still applies, access control, secured infrastructure, patched dependencies, and your existing security programme covers it. Two presumption routes can also do work here, and both save effort rather than adding it.

The first sits in the Act itself. A high-risk system certified, or holding a statement of conformity, under a European cybersecurity certification scheme adopted pursuant to the Cybersecurity Act is presumed to comply with this Article’s cybersecurity requirements, so far as the certificate covers them [1, Art 42(2)].

The second comes from next door. Where a high-risk system is also a product with digital elements in scope of the Cyber Resilience Act, meeting the CRA’s essential cybersecurity requirements means it is deemed to comply with the cybersecurity requirements of this Article, subject to the conditions in Article 12 of that Regulation [28, Art 12].

Note where that leaves the paperwork. Neither route discharges accuracy or robustness. Both are scoped to what the certificate or the CRA assessment actually covers, and your Annex IV Part Seven description must say which requirements you are relying on them for and which you are meeting some other way. One assessment feeding two regimes is a real saving. One assessment assumed to cover everything is a finding. Chapter 20 maps the wider neighbourhood including NIS2.

What Article 15(5) adds is the AI-specific attack surface, and the Act names the families itself. Data poisoning, meaning corruption of the training data. Model poisoning, meaning corruption of pre-trained components. Adversarial examples and model evasion, inputs crafted to make the model fail. Confidentiality attacks, extracting the model or its training data. And model flaws.

Each family points at a control, and the mapping is straightforward to write down. Poisoning attacks point at supply chain controls: provenance for training data, integrity checks on datasets, and vetting of every pre-trained model and third-party component you build on. 

Chapter 7’s data governance file already holds half of this, which is a pleasant surprise in a book not otherwise generous with them. Evasion attacks point at adversarial testing: someone on your side tries to fool the model before someone on the other side does, and the attempts and results go in the test file. 

Confidentiality attacks point at query controls: rate limiting, output filtering, and monitoring for the query patterns extraction requires. Model flaws point back at the whole of Part Two, disciplined development, documented and reviewed.

Then scale the effort to the incentive. Ask who profits from breaking this particular system, and by how much. A fraud model faces motivated, funded, iterating adversaries. Every fraudster it blocks is running experiments against it, daily, for money. 

An internal document classifier faces almost nobody, unless your filing system has enemies you have not told anyone about. The first system needs red-teaming on a schedule. The second needs the basics. 

Write that reasoning down, because it is your “appropriate level” justification for this limb, and it is exactly the kind of proportionality argument authorities accept when it was documented at the time and dismiss when it was retrofitted afterwards.

Testing regimes: the evidence factory

All three limbs converge on the same practical machinery, a testing regime that starts before market and never stops. The shape is the same for each limb.

Before market

A test plan per system: what gets tested, against which metrics, on which data, at which thresholds. Then execute it, with results recorded, dated and signed, because Annex IV Part Two asks for test logs and reports in exactly that form. 

The failures stay in the file, along with the fixes that followed them. A test file containing only passes is not a good sign. It is a sign the tests were written to pass, and examiners have seen enough test files to know the difference.

At every change

Chapter 10’s release gate carries the load here. A retrained model, a new data source, a dependency upgrade: each one triggers the affected tests again, and the results version alongside the documentation. 

The examiner’s question is brutal in its simplicity. This model version, the one in production, show me its test results. Not the previous version’s. This one’s.

In production, continuously

Drift monitoring against the declared accuracy. Anomaly detection on inputs and outputs. For adversarially exposed systems, monitoring for attack patterns. All of it flows into the Chapter 11 logs and surfaces in the Chapter 18 post-market reports.

Periodically, adversarially

For the systems where the incentive analysis says so: scheduled red-team exercises, scoped, dated, with findings tracked to closure. And note that the closure records matter more than the findings themselves. Every system has vulnerabilities. The file must show that yours get fixed.

What “appropriate” means when no standard answers you

Now the hard word. The Act’s design assumes harmonised standards will eventually define appropriate levels. Build to the standard, gain a presumption of conformity under Article 40, everyone goes home early. 

The standards, from CEN and CENELEC’s joint committee, are arriving slowly, and for many system types nothing final will exist even by the December 2027 date. The Commission is separately directed to encourage the development of benchmarks and measurement methodologies for these requirements [1, Art 15(2)], which is the legislative equivalent of asking someone to hurry.

Until the standards land, “appropriate” is a judgement you must construct and defend yourself. Annex IV Part Seven already gives you the format: where harmonised standards were not applied, describe the solutions adopted to meet the requirements instead. Here is how to build that description so it survives scrutiny.

Anchor to the risk file

Appropriate means appropriate to the harm. Start from the Article 9 risk assessment: these are the failure modes, these are their consequences, and therefore these are the accuracy floors, the robustness scenarios and the security controls that address them. A justification that starts from the risks is an argument. A justification that starts from whatever the system happens to achieve is an excuse wearing an argument’s clothes.

Borrow the nearest respectable benchmark

No harmonised standard does not mean no standards at all. ISO/IEC 42001 for the management layer. ISO/IEC 23894 for AI risk. ISO/IEC 24029 for robustness of neural networks. The NIST Adversarial Machine Learning taxonomy for the attack families. Sector norms wherever they exist. Citing recognised technical references shows your judgement was calibrated against the state of the art, and state of the art is the phrase the Act’s entire conformity logic leans on.

Compare against the alternative

Here is one argument that regulators and courts both understand instinctively: the system was tested against the process it replaced. If human reviewers achieved X, and the system achieves better than X on the same measure, with the comparison documented, then “appropriate” has a floor under it. If nobody ever measured the process being replaced, that is a fact worth discovering before an authority asks, rather than during.

Date your judgement and diarise its expiry

The state of the art moves, and the standards will eventually publish. A justification written in 2026 needs a review trigger for the day the relevant standard lands, because from that day the question changes. It stops being “was your judgement reasonable” and becomes “why are you not following the standard”. Chapter 5’s trigger discipline covers this. Make standards publication one of the watched events.

The honest summary of this whole section: in the standards gap, you cannot buy certainty. You can only build defensibility. Defensibility is a written chain running from risk to requirement to test to result, calibrated against the best available references, reviewed on triggers. 

That chain is buildable today, with tools you already have. Organisations waiting for the standards to make the judgement for them are confusing the arrival of a benchmark with the arrival of an excuse.

What an auditor will ask for

For this chapter, expect the following requests: 

  • The declared accuracy metrics for a sampled system, then the test reports that produced them, then the production monitoring showing the declaration still holds today. 
  • Per-group performance figures wherever the risk file identified specific populations. 
  • The robustness test plan and its results, including the failure behaviour and the manual fallback with its named owner. 
  • For learning systems, the feedback loop analysis. 
  • The AI-specific threat assessment, mapped to the Article 15(5) attack families, with the incentive reasoning attached. 
  • Red-team reports, and the closure records for their findings. 
  • The current model version’s test results, not last quarter’s. 
  • Wherever no harmonised standard was applied, the Annex IV Part Seven justification, complete with its references, its date, and its review trigger.

Article 15 closes the requirements the system itself must meet. The next chapter leaves the high-risk tier and turns to the transparency duties that arrive first and reach furthest: Article 50.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 14: Transparency for Limited-Risk Systems, Article 50

Everything in Part III so far concerned high-risk systems, with deadlines in December 2027 and beyond. This chapter is different, and the difference is the whole point. Article 50 applies from 2 August 2026. It reaches systems in every risk tier. 

And it is the first AI Act obligation most organisations will actually feel. The Digital Omnibus deferred the high-risk calendar and left this date standing, together with the enforcement powers that arrive alongside it [8].

The logic of Article 50 is simple enough to state in one breath. People should know when they are dealing with a machine, and when content is synthetic. The article turns that idea into four distinct duties, two sitting on providers and two on deployers. Most organisations owe at least one of them. 

Many owe several, and have not yet worked out which.

Duty one: tell people they are talking to a machine

Providers must design AI systems that interact directly with people so that those people are informed they are interacting with an AI system, unless that fact is obvious to a reasonably well-informed, observant and circumspect person, given the circumstances and the context of use [1, Art 50(1)].

Let’s unpack the working parts.

It is a design duty on the provider. The disclosure must be built into the system, not left for the deployer to remember. If you build or white-label a chatbot, a voice agent or any conversational interface, this obligation is yours. Chapter 2’s Article 25 warning applies with full force: put your brand on a procured chatbot, and the duty travels with the brand.

The trigger is direct interaction. Chatbots, voice assistants, phone agents, and the growing population of AI agents that message people, book things and answer on an organisation’s behalf. The more human the interaction feels, the harder the duty bites, and agentic products that hold long conversations are exactly what the drafting anticipates.

The exception is narrow, and it does not reward wishful thinking. “Obvious to a reasonably well-informed person” covers the video game character and the plainly labelled demo. It does not cover the customer service chat whose widget was designed to feel human, and it certainly does not cover a voice agent that answers the phone. The test is context, and the burden of judging the context wrongly falls on you. 

So here is the safe reading. If a designer worked to make the interaction feel human, disclose. The effort spent making it feel human is itself evidence that the machine-ness was not obvious. You cannot polish away the tell and then claim everyone saw it.

What compliance looks like in practice is almost embarrassingly cheap. A clear statement at the start of the interaction, in the language of the interaction, before the substance begins. “You are chatting with an AI assistant” costs nothing. For voice, it is one sentence at the top of the call. 

The information must reach the person at the first interaction or exposure at the latest, and it must be clear and distinguishable [1, Art 50(5)].

Duty two: mark synthetic content at the source

Providers of AI systems that generate synthetic audio, image, video or text must ensure the outputs are marked in a machine-readable format and detectable as artificially generated or manipulated [1, Art 50(2)]. 

The solutions must be effective, interoperable, robust and reliable, as far as technically feasible, taking account of the state of the art and the limitations of different content types.

This is the watermarking duty, and it sits on the provider of the generative system, not on its users. If your product generates images, video, audio or text, you owe outputs that a detection tool can identify as synthetic.

The Omnibus adjusted the timing, and it did so precisely. Systems already on the market before 2 August 2026 have until 2 December 2026 to comply [4]. Systems placed on the market from 2 August 2026 onwards must comply from the day they launch. 

Read that twice if you are launching a generative product this autumn. The four-month reprieve belongs to your established competitors. Not to you.

“Machine-readable” is the phrase doing the heavy lifting, and the marking-in-practice section below deals with it properly. 

One exception is worth knowing now. Systems performing an assistive function for standard editing, or that do not substantially alter the deployer’s input data or its meaning, sit outside the duty. Your spell-checker and your photo touch-up tool are not deepfake factories, and the Act says so.

Duty three: disclose emotion recognition and biometric categorisation

Deployers of emotion recognition systems or biometric categorisation systems must inform the persons exposed to them about the operation of the system [1, Art 50(3)].

A short duty with sharp edges. Note who owes it: the deployer, the organisation actually running the system on people, not the vendor who sold it. And note what Chapter 4 already established. Emotion recognition in workplaces and educational institutions is prohibited outright, so this disclosure duty governs whatever deployments remain lawful, with customer-facing analytics as the main surviving category.

If your organisation runs anything that infers emotional states or sorts people by biometric characteristics, two questions apply, in strict order. Is it prohibited? And if not, have the exposed people been told? A retail analytics deployment that fails the second question is one complaint away from an authority asking the first.

Duty four: label deepfakes and synthetic public-interest text

If you publish deepfakes, you say so. Deployers of AI systems that generate or manipulate image, audio or video content constituting a deepfake must disclose that the content is artificially generated or manipulated [1, Art 50(4)].

A deepfake, in the Act’s definition, is content resembling real persons, objects, places, entities or events that a person would falsely take for authentic [1, Art 3(60)]. In other words, if it looks real and is not, it qualifies.

This is the duty that lands on marketing departments, most of whom do not yet know they have it. The provider marked the content invisibly under duty two. 

You, the publisher, owe the visible label. The synthetic presenter in your advert, who has never existed and therefore works weekends. The product photos of a warehouse you do not own. The founder’s voice in the podcast intro, cloned because the founder was busy that week. If a reasonable viewer could take it for real, label it [1, Art 50(4)].

Two escape routes exist, and both are narrower than people hope.

The first covers art. Where the content forms part of an evidently artistic, creative, satirical or fictional work, the duty shrinks to a disclosure that does not spoil the work. A credit line, not a stamp across the frame. Note the word evidently. Your product advert is not satire because the legal team says so after the complaint arrives.

The second covers text. The labelling duty applies to AI-generated text published to inform the public on matters of public interest, unless a human reviewed it and someone holds editorial responsibility. 

For ordinary corporate use this resolves neatly. AI-drafted, human-reviewed, editorially owned means no label required. AI-generated news pushed out with no human review means label it, though an organisation publishing unreviewed machine journalism has problems no label will fix. The label just tells everyone what they are reading.

Machine-readable marking in practice

The Act says synthetic content must be marked in a “machine-readable format”. That phrase needs translating into decisions, because several technologies claim to answer it, and none answers it completely.

The candidates, in order of maturity.

Metadata

Provenance information written into the file itself, with C2PA content credentials as the emerging reference standard. The file carries a cryptographically signed note saying what it is and where it came from. Easy to implement. Also easy to remove, since a screenshot defeats it, which is a modest bar for an adversary to clear.

Invisible watermarking

A signal embedded in the pixels or the audio, designed to survive editing and re-encoding, readable by detection tools. Sturdier than metadata, but locked in a permanent arms race with the people working on removal. Both sides are well funded.

Visible marks

Labels burned into the content. These serve duty four’s human audience rather than duty two’s machines, so they complement the technologies above without replacing them.

Text

The problem child of the family. Statistical watermarking of word choices exists, and it survives until someone paraphrases the text, which is to say it survives until someone tries. Text marking remains the least settled corner of the field, and everyone working in it knows.

So nothing works perfectly, and the Act knows it. The duty applies “as far as technically feasible, taking account of the state of the art”, which is the legislator saying: we are aware the tools are imperfect, use the best available combination anyway, and keep pace as they improve. 

You are not required to achieve the impossible. You are required to stop pretending the possible is impossible.

For a generative provider in 2026, a defensible implementation looks like this. Metadata provenance on every output. Invisible watermarking wherever the medium is mature, which today means images and audio, with text trailing behind. 

A written record of the choices, the state-of-the-art assessment behind them, and a review trigger for when the AI Office guidance lands. And a working detection route, because the duty says detectable, and detectable-in-principle is not a category the Act recognises. File the whole package under Chapter 10 disciplines, versioned with the product.

For everyone else, this is a procurement question. If you buy generative capability and publish its outputs, the deepfake labelling duty is yours, and you can only label what you know to be synthetic. A vendor whose outputs carry no provenance is selling you a labelling obligation with no way to meet it, which is generous of them. Add marking to the procurement checklist from Chapter 9.

And finally, check your own plumbing. Most image pipelines strip metadata by default, as an optimisation. Which means a content workflow can take a properly marked file, helpfully delete the mark, and publish the result, leaving you non-compliant through the quiet diligence of your own infrastructure. Provenance has to survive your systems, not merely arrive at them.

Getting ready: the six-week version

For an organisation reading this against the August 2026 date, the sequence is short.

Start with the inventory. Map the four duties against the Chapter 8 system list. Which systems interact with people directly. Which generate synthetic content. Whether anything touches emotion recognition or biometric categorisation. And where AI-generated content flows into your publishing.

Fix the interaction disclosures first. They are the cheapest to implement, one sentence in an interface, and the most visible when absent.

Then the contracts. For procured systems, confirm in writing which side implements which duty. Provider and deployer duties interlock throughout this article, and the gap between two organisations each assuming the other has it covered is exactly where the first enforcement cases will be found.

Then publishing. A labelling rule for synthetic media in the content workflow, with the artistic-work credit-line variant defined, and provenance preservation verified in the pipeline.

And finally the evidence. A one-page record per duty. What was implemented, when, and on whose decision, following the pattern this book has used throughout.

What an auditor will ask for

For this chapter, expect the following requests. 

  • The inventory mapping of the four duties across your systems. 
  • For interactive systems, the disclosure exactly as the user sees it, in each supported language, plus the design record showing it appears at first interaction. 
  • For generative products, the marking implementation, its state-of-the-art justification, its detection route, and the version history against the December 2026 date or your placing-on-market date, whichever applied. 
  • For emotion or biometric deployments, the Chapter 4 prohibition analysis first, then the disclosure text and where exposed persons actually encounter it. 
  • For published synthetic content, the labelling rule, a sampled published item with its label, and the editorial-responsibility record wherever the text exemption was relied on. 
  • For procured systems, the contractual allocation of every one of these duties.

Article 50 is where the Act meets the public. The next chapter turns to the layer beneath every generative product: the obligations of general-purpose model providers, and what the rest of the value chain is entitled to demand from them.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 15: GPAI Obligations in the Value Chain

Almost nothing in modern AI is built from scratch. The chatbot on your website, the drafting assistant in your legal team, the classifier inside your product: underneath most of them sits a general-purpose model built by someone else, trained at a cost you do not want to think about, and licensed to you through an API and a prayer.

The Act noticed. Chapter V regulates the models themselves, separately from the systems built on them, and it has been in application since 2 August 2025, which is longer than most of the obligations in this book. From 2 August 2026 the Commission can fine model providers for breaching it [8], which converts the regime from an honour system into law with consequences.

This chapter covers what model providers owe, what flows down the chain to you, when you accidentally become a model provider yourself, and how to read a model provider’s documentation the way a due-diligence lawyer reads a data room. Which is to say, looking for what is missing.

What a model provider owes

Article 53 gives every provider of a general-purpose AI model four duties [1, Art 53(1)].

Technical documentation of the model, following Annex XI. The training and testing process, evaluation results, architecture, compute, energy consumption. Kept up to date, and available to the AI Office and national authorities on request.

Documentation for downstream providers, following Annex XII. This is the information a builder needs to integrate the model into a system and meet their own obligations. Capabilities, limitations, how to use the thing properly. Of the four duties, this is the one that exists for your benefit, and we return to it below.

A copyright policy. A policy to comply with EU copyright law, including identifying and respecting the rights reservations that rightsholders express under the text and data mining rules. In plain terms, a written answer to the question: what did you do about the people who said do not train on my work?

A public training-content summary. A sufficiently detailed summary of the content used to train the model, published using the AI Office’s template, which came out in July 2025 [22]. This is the most publicly visible of the four duties, and the one most closely watched by rightsholders with lawyers.

One large exemption sits alongside. Providers of models released under a free and open-source licence, with weights, architecture and usage information made publicly available, are exempt from the first two duties, the documentation ones. They still owe the copyright policy and the training summary. 

The exemption evaporates entirely if the model poses systemic risk, on which more shortly. So “it’s open source” answers some questions and not others. Keep that sentence ready for vendor meetings, because you will need it.

A legacy note to close the section. Models placed on the market before 2 August 2025 have until 2 August 2027 to comply [9]. If a model provider’s file looks thin and their model is old enough, that may be why. It is an explanation. It is not a comfort.

Systemic-risk models: the heavier regime

Some models are large enough that the Act treats them as infrastructure with opinions. A GPAI model poses systemic risk by exactly two routes [1, Art 51]. Either it has high-impact capabilities, which are presumed when cumulative training compute exceeds 10²⁵ floating-point operations. Or the Commission designates it, using the Annex XIII criteria, capabilities, reach, user numbers and so on.

Note what is not on that list, because vendor conversations regularly invent additions to it. Open weights do not trigger systemic risk. User counts alone do not trigger it, though they feed the Commission’s designation criteria. There is no AI Office safety score. The routes are compute and designation, nothing else. 

Providers crossing the compute line must notify the Commission themselves, within two weeks, which is the regulatory equivalent of being required to raise your own hand.

Systemic-risk providers owe everything in Article 53 plus the Article 55 layer [1, Art 55]: 

  • Model evaluations including adversarial testing. 
  • Assessment and mitigation of systemic risks at Union level. 
  • Serious incident tracking and reporting to the AI Office. 
  • Adequate cybersecurity for the model and the physical infrastructure it lives on.

For most readers the practical relevance is indirect, but it is real. If your product is built on a frontier model, your supplier is running this regime, or should be, and their compliance posture is part of your supply chain risk. The frontier labs’ safety frameworks and evaluation reports are not marketing exercises. They are Article 55 evidence, and you are allowed to read them that way.

The Code of Practice

The Act invited the industry to write down how it would comply, and the result is the General-Purpose AI Code of Practice, published in July 2025, with chapters on transparency, copyright, and safety and security [23].

Signing is voluntary. Most major model providers signed anyway, because the alternative to demonstrating compliance through the Code is demonstrating it some other way, alone, to a regulator with questions. Few volunteered for that.

For downstream businesses the Code has one practical use. It is a checklist of what good looks like. A model provider who signed has committed to specific transparency and copyright practices, including a model documentation form that conveniently packages the Annex XII information into one place. A provider who did not sign is entitled to that choice, and should be asked, politely, what they are doing instead. The answer will be informative either way.

When you become a model provider by accident

Chapter 2 planted this warning, and here comes the payoff. If you take a general-purpose model and fine-tune or otherwise modify it, you may become the provider of a new model, with Article 53 duties attaching to your modification. 

The Commission’s July 2025 guidelines draw the line with a compute threshold: the modification becomes model provision when the compute used for it exceeds roughly one third of the original model’s training compute [13].

For ordinary corporate fine-tuning, a few thousand examples to teach a model your house style, the threshold is comfortably distant. For AI companies doing serious continued pre-training, it is a live question, and it belongs in the training plan, costed, before the run starts. 

The analysis takes one page. The retrofit takes a compliance programme. The choice between them takes place exactly once, before the training run, and never again.

The Omnibus added a wrinkle at the other end of the chain: supervision. Where a provider builds systems on its own GPAI model, the AI Office’s supervisory role over those systems was strengthened [4].

Reading a model provider’s documentation like a due-diligence lawyer

Now the skill this chapter exists to teach. When you build on a general-purpose model, or buy systems built on one, the model provider’s documentation is your evidence foundation. And due diligence on it is not reading what is there. It is noticing what is not.

Start with the Annex XII package, the one owed to you:

  • Does it exist at all, as a coherent document? Or is it a documentation website plus a blog post plus a shrug? 
  • Does it state the model’s capabilities and limitations in terms you could put in front of your own Chapter 12 oversight design? 
  • Does it tell you enough about intended and excluded uses to know whether your use case sits inside the fence? 

And watch for one particular pattern. A provider whose acceptable use policy is more detailed than their limitations documentation has told you what they fear. It is not your compliance.

Then the training-content summary. It is public and it is mandatory, so find it. Read it against the AI Office template and note the level of generality. Then read it next to your own exposure. If your business publishes content, is your content plausibly in there? And what did the copyright policy say about opt-outs? 

If your legal team has strong feelings about your material appearing in training data, this document is where the feelings meet the facts.

Then the version question, the one that catches everyone. Which model version does the documentation describe, and which version does your API actually serve today? Model providers update silently and often. 

Your Chapter 10 documentation, your Chapter 11 logs and your Chapter 13 test results are all keyed to model versions, and a provider who swaps the model underneath your integration has just invalidated evidence you are legally required to keep current. 

The contract needs version pinning or change notice. Your logging needs the model version on every transaction, which Chapter 11 already told you to capture. Now you know the second reason.

Then the systemic-risk posture, where it applies. Is the model over the compute threshold, or designated? If so, where is the safety framework, where are the evaluation reports, where is the incident channel? These are Article 55 artefacts. Their absence from a frontier provider is a finding, not a gap.

And finally, the contractual mesh. The Act creates duties along the value chain, but the contract decides how they operate between the two of you. What flows down: documentation, change notices, incident notifications, and cooperation when your authority asks you questions the model provider must help answer. What flows up: your usage data, your fine-tuning data, and, in some standard terms, rather more than you intended. 

Read the training-use clause twice. Some vendors reserve the right to learn from your data, which means your confidential inputs may become tomorrow’s training content, summarised in public under Article 53(1)(d). That clause deserves better than a skim, and it rarely gets it.

Record the whole exercise on one page per model dependency. Provider, model, version discipline, documentation received, gaps found, questions asked, answers filed. This becomes the model-layer entry in your evidence map. When a customer’s due-diligence team asks how you govern your AI supply chain, it is the page you show them.

What you owe downstream

The chain does not end with you. If you build systems on GPAI models and sell them, you are the downstream provider today and someone else’s upstream tomorrow. Everything you just demanded, your customers are entitled to demand from you. Honest capability and limitation documentation. Version discipline. Change notices. Answers that arrive before their regulator’s deadline rather than after it.

The value chain is a mirror. Build your own Annex XII-equivalent pack once, keep it versioned, and hand it over without being chased. If the principle does not move you, Chapter 22 is full of organisations that did the opposite, and none of them enjoyed how the story ended.

What an auditor will ask for

For this chapter, expect the following requests:

  • Your model dependency register, listing every GPAI model your systems rely on, directly or through vendors. 
  • Per dependency, the due-diligence record described above: the Annex XII documentation received, the training-content summary reviewed, and the contract clauses on versioning, change notice and data use. 
  • Your fine-tuning analysis against the compute threshold, dated before the training run it assesses, because a dated-after analysis is a different document with a different meaning. 
  • For anything you supply downstream, your own documentation pack, plus evidence that you actually deliver change notices. 
  • Per transaction in your logs, the model version that produced it. That single field is what joins your entire evidence chain to a supplier’s silent update.

The model layer is where your compliance rests on someone else’s engineering. The next part of the book turns to the machinery that judges the whole construction: conformity assessment, audit, and what happens when something goes wrong.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Part IV — Auditing and Enforcement

Parts II and III built the compliance machine. Governance, documentation, logs, oversight, testing. Part IV is about the moments when someone outside the organisation examines it.

Four such moments exist. Conformity assessment, when the system is judged before it reaches the market. Inspection, when an authority arrives with questions and legal powers. Incidents, when something goes wrong in public and the reporting clocks start ticking. And enforcement, when the examination ends badly and the fine schedule becomes personally interesting.

The chapters in this part share one premise. None of these moments can be survived through improvisation. Every one of them is won or lost months earlier, in the quality of the artefacts the earlier chapters built. 

What Part IV adds is the machinery of the encounters themselves. Who examines, under what powers, against what checklist, on what timeline, and what they take away with them. Read it before you need it. The organisations that read it afterwards are the case studies in Chapter 22.

Chapter 16: Conformity Assessment

Before a high-risk AI system may be placed on the EU market, someone must formally conclude that it complies with the requirements of Chapter III Section 2, which is to say, everything Part III of this book just built. That conclusion has a name, conformity assessment. It has a paper trail. And it has a surprise: for most high-risk AI systems, the person who performs it is you.

This chapter covers the two assessment routes, the declaration and marking that follow, the database registration, and the part that deserves far more respect than it usually gets. What you are legally signing when you sign.

The two routes

Article 43 provides two procedures [1, Art 43].

Internal control, under Annex VI

The provider verifies its own quality management system, examines its own technical documentation, and confirms its own system’s conformity. There is no external examiner anywhere in the process. The provider assesses, the provider concludes, the provider signs.

Notified body assessment, under Annex VII

An independent conformity assessment body, notified to the Commission by a Member State, examines the quality management system and the technical documentation, then issues a certificate. External examiner, external conclusion, and an external party whose name is now attached to your product.

The allocation between the two routes is the surprise. For the Annex III high-risk systems, employment, credit, education, essential services, the rest of Chapter 5’s list, the route is internal control. 

Only one Annex III area can require a notified body at all, and that is biometrics, and even there only where harmonised standards or common specifications are unavailable or not fully applied [1, Art 43(1)]. 

For Annex I systems, meaning AI inside regulated products, conformity assessment follows the product regime it inherits. Your medical device’s notified body absorbs the AI requirements into its existing examination [1, Art 43(3)]. The final Omnibus text keeps that route moving through the transition: bodies already notified under the Annex I legislation may assess AI conformity for eighteen months from 27 July 2026, applying for full designation under this Regulation by 28 January 2028 [1, Art 43(3)].

So the recruitment screener, the credit model and the exam proctor are all self-assessed. The legislator looked at the volume of systems, looked at the number of notified bodies in existence, and made a pragmatic choice with a sharp edge. The examination is internal, but the conclusion is public and binding. The Act did not lower the bar. It handed you the clipboard.

And assessment is not once-and-done. A substantial modification, Chapter 2’s term, triggers a fresh conformity assessment for the modified system [1, Art 43(4)]. There is one carve-out, built for learning systems. 

Changes to the system and its performance that the provider pre-determined at the moment of the initial assessment, and documented in the technical documentation, do not count as substantial modifications. 

Which means the pre-determined changes section of Annex IV, the one that looked like bureaucratic filler back in Chapter 10, turns out to be the clause that lets your model keep learning without re-certification. Draft it with care and ambition.

Harmonised standards: the presumption you want

Systems built to harmonised standards published in the Official Journal are presumed to conform with the requirements those standards cover [1, Art 40]. The presumption converts an argument into a citation. “Our approach is appropriate” becomes “we applied EN standard X”, and the second sentence is much easier to defend.

Chapter 13 covered the current shortage of such standards, and the self-built justification that fills the gap. The consequence for conformity assessment is that early assessments will lean heavily on the Annex IV Part Seven descriptions, and the file must carry the full reasoning the presumption would otherwise have replaced. 

When the standards eventually land, reassessment against them becomes the cheaper path. One more entry for the trigger list.

The declaration of conformity

The assessment concludes in a document, the EU declaration of conformity. The provider draws one up for each high-risk system. It states that the system meets the Section 2 requirements, identifies the system, names the standards or specifications applied, gets kept for ten years, and goes to authorities on request. By drawing it up, says the Act with unusual directness, the provider assumes responsibility for compliance with the requirements [1, Art 47].

Where other Union law also requires a declaration, one combined declaration covers both regimes, which spares the medical device manufacturer a second ceremony.

The declaration is one page long. Treat the drafting with the seriousness that brevity invites people to skip. The system identification must match the technical documentation. The versions must match reality. The standards list must match what was actually applied, rather than what was intended back when everyone was optimistic. 

A declaration citing a standard the engineering team abandoned in March is a false statement with a signature on it.

CE marking and registration

Two visible consequences follow the declaration.

The CE marking goes onto the system, affixed visibly, legibly and indelibly, or for digital systems, made accessible through the interface or the accompanying documentation [1, Art 48]. The familiar two letters from toasters and toys now appear on software, and they mean exactly what they have always meant. The manufacturer declares conformity with applicable Union law. Where a notified body was involved, its identification number sits alongside the marking.

Then registration. Before placing the system on the market, the provider registers itself and the system in the EU database, the public register the Commission maintains under Article 71 [1, Arts 49, 71]. Deployers who are public authorities register their use as well. Understand what the database entry is: public-facing compliance. Your system and your declaration, visible to authorities, competitors, journalists and claimants, all with equal convenience.

One registration point needs stating carefully, because outdated commentary still circulates. The original Act also required providers relying on the Article 6(3) filter, our system sits in an Annex III area but is not high-risk, to register those systems. 

The Commission’s Omnibus proposal deleted that duty. The data protection authorities objected, the Parliament and Council both refused the deletion, and the final text keeps the registration duty in a simplified form, with a trimmed public entry [1, Art 49(2)] [21]. 

Notably, the summary of the exemption grounds no longer appears in the public register, which makes your internal Article 6(4) assessment the only place your reasoning lives. Chapter 5 drew the conclusion, and it bears repeating here. The private record became more load-bearing, not less.

Self-certification as a legal act

Now the section this chapter exists for. Internal control feels administrative. Your own team, your own documents, your own conclusion, all in-house and unhurried. The feeling is wrong, and the wrongness has structure.

The declaration is a representation with statutory consequences. Article 47 says the words plainly: drawing it up means assuming responsibility. And supplying incorrect, incomplete or misleading information to notified bodies and authorities carries its own fine tier [1, Art 99]. 

A declaration that overstates conformity is not an optimistic form. It is the document an enforcement case gets built around, because it proves two things at once. You knew the requirements, and you stated you met them.

It radiates into private law. The declaration and the CE marking are public statements of conformity. In a negligence claim after the system harms someone. In a contract dispute with a deployer who relied on the marking. Under the revised Product Liability Directive, which reaches software and AI. In each of these, the declaration is the claimant’s first exhibit. The defendant certified this.

Your customers will make it a warranty too. Enterprise contracts increasingly copy the declaration’s substance into representations with indemnities attached, which converts a regulatory statement into a commercial one, with a price.

In addition, it has a signatory. A natural person signs for the provider, and the organisation must decide who that person is. The honest answer runs: someone senior enough to bear it, informed enough to mean it, and supported by a file that lets them mean it honestly. 

Which produces, at last, the practical machinery this whole book has been building: 

  • The declaration should be the top page of a signing pack. 
  • The technical documentation current. 
  • The test results keyed to the shipping version. 
  • The risk file reviewed. 
  • The open findings listed, with a named acceptance beside each one. 

If the pack cannot support the sentence “this system complies”, the pack is telling you something, and the time to hear it is before signing, not after.

Here is the comparison that concentrates minds. Internal control under the AI Act works like a directors’ solvency statement. Nobody external checks it at the moment of signing. Everybody external checks it the moment something fails. The absence of an examiner at the start is not the absence of examination. It is the deferral of examination to the worst possible moment, with your signature waiting there when it comes.

Deployers, one paragraph for you. You will receive declarations and see CE markings, and your Chapter 1 verification instinct should engage on contact. The declaration tells you the provider got certified. It does not tell you the certification was sound.

For consequential systems, ask for the substance behind it, the instructions for use you were owed anyway, the accuracy declaration, the oversight design. A provider offended by the question has answered it.

What an auditor will ask for

For this chapter, expect the following requests. 

  • The conformity assessment record for a sampled system, meaning the Annex VI procedure as it was actually performed, dated, with the quality management system verification and the documentation examination it claims to contain. 
  • The declaration of conformity, checked against the technical documentation, against the shipping version, and against the standards actually applied. 
  • The CE marking as a user encounters it. The database registration, current, and matching the declaration. 
  • For modified systems, the substantial-modification analysis and the reassessment, or the pre-determined-changes clause that excused it. 
  • The signing pack behind the declaration, together with the name, role and briefing of the person who signed. 
  • For anything assessed against self-built justifications rather than standards, the Part Seven reasoning, plus the trigger entry showing you will reassess when the standard publishes.

The assessment happens before the market. The next chapter deals with what happens after. The audit-ready organisation, and the day the examiner is real.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 17: The Audit-Ready Organisation

Every chapter so far has ended with the same section, what an auditor will ask for. This chapter is about the day the auditor stops being hypothetical. It covers who can examine you and with what powers, how inspections actually unfold, the one artefact that holds the whole file together, and a method for finding your gaps before someone with enforcement powers finds them for you.

One reframe before the machinery. Audit-readiness is not a special state you enter when a letter arrives. It is a property your compliance programme either has or lacks on any given Tuesday, and it reduces to a single question. 

Can you produce the evidence, quickly, in a form an outsider can follow, without heroics? 

Organisations that need three weeks and a war room to answer a document request are not audit-ready. They are audit-adjacent, and the difference is perfectly visible from the regulator’s side of the table.

Who examines you, and with what powers

Enforcement of the AI Act runs through market surveillance, the same machinery that polices toys, machinery and medical devices, applied under Regulation (EU) 2019/1020 [1, Art 74] [27]. Each Member State designates its market surveillance authorities, and several have spread the work across sectoral regulators. 

So your examiner may be the financial supervisor, the data protection authority, or a labour inspectorate wearing an additional hat, depending on the system and the country. The AI Office supervises GPAI models, and the Omnibus strengthened its role over systems a provider builds on its own model [4].

The powers are the point, and they are considerable. Under the market surveillance framework, authorities can require documents and information, conduct on-site inspections, acquire and test product samples, and, where non-compliance persists, order corrective action, restrict or prohibit the making available of a system, and order its withdrawal or recall.

The AI Act then adds the specific keys. Full access to the technical documentation. Access to the automatically generated logs under the provider’s or deployer’s control. Where necessary to assess conformity, and on reasoned request, access to the training, validation and testing datasets. And for high-risk systems, access to source code itself, where specified conditions are met.

Two implications deserve a sentence each. First, everything Part III told you to build is legally reachable, the documentation, the logs, the test files, and in defined circumstances the code. Second, the reasoned-request structure makes the encounter iterative. Authorities ask, escalate, and ask again, which means your first response sets the tone for everything that follows it.

How inspections actually run

Inspections vary by country and by trigger, but the anatomy is stable enough to prepare against. Five phases.

The trigger

Most inspections are not random. They start from a complaint, an affected person, a competitor, an NGO. Or from an incident report you filed under Article 73. Or from a referral by another authority, with the data protection authority as the classic feeder. Or from media coverage. Or from a sector sweep, where an authority examines everyone in a category at once. You will rarely know the trigger at the start, which is worth remembering while drafting responses. The authority may know considerably more than the first letter reveals.

The letter

A written request for information and documents, with a deadline, typically weeks rather than months. This is the phase where audit-readiness gets measured in the most literal way, because the requests track the obligations, and the obligations are the ones this book has been mapping chapter by chapter: 

  • Classification records. 
  • The technical documentation. 
  • The declaration. 
  • Logs for a stated period. 
  • The oversight design and its metrics.

The exchange

Follow-up questions, progressively narrowing. What the authority is doing in this phase is triangulating. Your documents against each other. Your declared accuracy against your logs. Your instructions for use against your marketing. Your stated oversight against the override rates. 

Chapter 10’s divergence review exists because this phase does, performed professionally, and with statutory powers behind it.

The visit, where one occurs

On-site or, increasingly, on-screen. Demonstrations of the system. Interviews with the people named in your records. The oversight interface shown live, the stop function actually performed. The people interviewed will be the named overseers and owners from your artefacts, which is why names in records should belong to people who know they are named, and know what the record says they do.

The finding

Compliance confirmed, or corrective action ordered with a deadline. For the serious cases, restriction, withdrawal, recall, and the penalty process of Chapter 19. Between those extremes sits the most common outcome of all, a negotiated remediation with a follow-up. And your credibility from phases two to four largely determines how much benefit of the doubt exists to negotiate with.

Three rules of conduct run through all five phases. 

  1. Respond on time, or ask for extensions early and with reasons, because missed deadlines convert document requests into findings. 
  2. Never supply a document that contradicts another document you have already supplied without flagging and explaining the difference yourself, because discovered contradictions read as concealment, while disclosed ones read as candour. 
  3. Treat every response as permanent. Authorities keep files, and the answer you give in a minor 2026 inquiry will be sitting in the folder your 2028 inspection opens with.

The evidence map

Everything above assumes you can actually find your evidence. The artefact that guarantees it is the evidence map, and if this book has a single deliverable, this is it.

The structure has four columns: 

  • Obligation: the specific duty, at article level, as it applies to a specific system. 
  • Artefact: the document, record or data that evidences compliance with it. 
  • Owner: the named person who maintains that artefact. 
  • Location: where it physically lives, the system and the path, precise enough that a stranger with access rights could retrieve it in minutes.

One row, for concreteness: Article 26(2) oversight competence, for the credit decisioning system. 

  • Artefact: the overseer assignment record and training log. 
  • Owner: the head of credit operations. 
  • Location: the compliance repository, oversight folder, current version. 

Multiply that by every applicable obligation for every in-scope system, and the map is complete. 

For a deployer with a modest inventory, that means a few hundred rows and roughly a fortnight’s work. Most of the fortnight goes on discovering where things actually are, which is, quietly, the point of the exercise.

The map earns its keep four ways: 

  1. It converts an authority’s letter into a retrieval exercise instead of an archaeology project. 
  2. It exposes gaps while they are still cheap, because an obligation row with no artefact next to it is a finding you made about yourself, privately. 
  3. It survives people, since owners leave, and a map with names and locations is the difference between succession and amnesia. 
  4. It compounds commercially, because the same map answers enterprise due diligence, insurer questionnaires and acquirer data rooms, which means the compliance function can, for once in its life, show the revenue side a reusable asset.

Maintain it like the living documentation of Chapter 10. Owned, versioned, reviewed on the same triggers, and reconciled to the inventory, so that every system’s obligations are mapped and every mapped artefact exists. 

An evidence map that has drifted from reality is more dangerous than no map at all. You will hand it over, and the authority will walk it, row by row.

Mock audit methodology

A mock audit is a rehearsal with consequences you control. The method takes two to four weeks of elapsed time, a lead with genuine independence from the artefacts being tested, your own counsel, an external reviewer, or an internal audit function, and the evidence map as its script. Six steps.

  1. Scope

Pick one system, your riskiest or most commercially exposed, and one role, provider or deployer. Resist the urge to audit everything. Depth finds gaps. Breadth finds nothing.

  1. Simulate the letter

Draft the document request a market surveillance authority would send for that system, using the chapter-end lists throughout this book as the source. Fifteen to twenty-five items. Date it, and start the clock you would face in reality.

  1. Retrieve cold

The team produces the documents using only the evidence map, without any help from the authors. Time it, and log every item that needs a person rather than the map to find. Retrieval time and author-dependence are your two readiness metrics, and the first mock audit’s numbers are usually educational in the way cold water is educational.

4. Triangulate. 

Now do what the authority does in the exchange phase, and read the retrieved set against itself. 

  • Intended purpose across the classification record, the technical documentation, the instructions for use and the marketing. 
  • Declared accuracy against the test reports and the production monitoring. 
  • Named people against the current staff list. 
  • Versions against what production actually runs. 

Log a finding for every mismatch, however small, because the examiner’s whole craft is small mismatches.

5. Interview. 

An hour each with two or three named individuals, the declaration signatory, an overseer, an artefact owner, asking what the records say they do and watching whether the answers match. This step is uncomfortable and non-optional. Documents that outrun their people fail in phase four of a real inspection.

6. Report and remediate. 

Findings graded: missing artefact, stale artefact, contradiction, undependable retrieval, person-record gap. Each with an owner and a date, tracked to closure like Chapter 13’s red-team findings, with the closure evidence filed. 

Then diarise the next mock audit, because the second one, a year later, is where you learn whether the programme self-heals or merely responds to prodding.

A worked walkthrough

Let us watch one happen. The firm is a fictional, mid-sized financial services company, but every finding below is the kind that turns up in real life, usually several at a time. The system under audit is a credit decisioning tool bought from a vendor.

Day one. The simulated authority letter arrives asking for nineteen items.

The first two hours go well. Twelve items come straight out of the evidence map. Then progress stops.

Four more items surface over the next two days. One was on a shared drive nobody had mapped. One was in the personal folder of an employee who left the company. Two were sitting in the vendor’s portal, which nobody had thought of as a place where compliance documents live. 

All four exist, so no harm done, except that the map said they were somewhere else, and in a real inspection, two days of searching is two days the authority notices.

Three items cannot be produced at all.

The fundamental rights impact assessment, which everyone believed was done, turns out to be a half-finished draft. Believed done is not done.

The log custody clause does not exist. The vendor contract was signed before anyone was thinking about logs, and it simply says nothing about who keeps them, for how long, or how the firm gets them out.

The instructions for use are stale. The vendor updated the model in the spring, and the firm is still holding the autumn edition, which describes a system that no longer exists.

Then the audit team starts reading documents against each other, and things get interesting.

The classification record says the system merely recommends, and a human decides. That was the whole basis for treating it as lower risk. But the operating procedure tells advisers to follow the score unless they can document an exception. 

Both statements cannot be true. A score you must follow by default is not a recommendation. It is the decision, with a human standing next to it for appearances. 

Nobody was careless here. The procedure simply drifted over time, quietly turning an advisory tool into a deciding one, and nobody thought to re-run the classification. An examiner who finds these two documents has found the whole case in two exhibits.

The numbers make it worse. Advisers override the score 0.4 per cent of the time, which is the rubber-stamping signature Chapter 12 warned about. The named overseer in the assignment record left in January. Her successor is competent and was trained, but none of it is written down anywhere. 

The retention schedule promises six months of logs, while the vendor portal exports ninety days. A promise the infrastructure cannot keep.

The interviews produce the finding that matters most. The vendor contract contains a compliance representation, the kind that quietly mirrors a declaration of conformity. It was signed by a procurement lead. Asked in his interview what the accuracy declaration meant, he answered honestly. He had assumed it was standard wording.

Fourteen findings. None of them exotic. Every one of them is cheaper to fix today than to explain to a regulator later.

The remediation plan takes eight weeks:

  • Finish the FRIA.
  • Amend the vendor contract to cover log custody, export and change notice. Write up the successor overseer properly. 
  • Either redraft the procedure to match the classification, or accept what the procedure reveals and reclassify the system. 
  • Correct the map with the four true locations.

The second mock audit is booked for the following spring, to test whether the fixes held.

That is the entire method. Nothing in it requires talent. It requires a calendar, and a tolerance for finding out.

What an auditor will ask for

For this chapter, the recursion is intentional. The auditor will ask for the evidence that you audit yourself. 

  • The evidence map, current and reconciled to the inventory. 
  • The mock audit reports, with findings, owners and closure evidence. 
  • The retrieval metrics, improving between cycles. 
  • The remediation of the walkthrough-class findings, contracts amended, successors recorded, contradictions resolved. 
  • The diary entry for the next cycle, because a single mock audit proves curiosity, while a cadence proves control.

The inspection tests the file you built. The next chapter covers the events that most often summon the inspector: post-market surveillance, incidents, and the reporting clocks of Articles 72 and 73.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 18: Post-Market Surveillance, Incidents and Corrective Action, Articles 72 and 73

Conformity assessment ends with a signature, a marking, and a warm sense of completion. Enjoy the feeling for a moment. The Act certainly does not intend it to last.

The moment your system reaches the market, a second regime begins. Watching the system in the field. Reporting when it hurts someone. Fixing or withdrawing it when the watching and the reporting reveal a problem. The launch was not the finish line. It was the start of the part with witnesses.

The premise deserves stating plainly, because it is the one providers most often miss.

  • Compliance at launch is a snapshot. 
  • Your accuracy declaration, your risk assessment, your oversight design, all of them are claims about how the system behaves. 
  • The field is where those claims meet reality, at scale, on actual people, none of whom read your test plan. 
  • Post-market surveillance is the mechanism by which you find out you were wrong before the newspapers do. 

That is genuinely the pitch. The Act is offering you the chance to be the first to know.

The monitoring system: Article 72

Providers of high-risk systems must establish a post-market monitoring system, proportionate to the technology and the risks, that actively and systematically collects, documents and analyses data on the system’s performance throughout its lifetime, and allows the provider to evaluate continuous compliance with the Chapter III requirements [1, Art 72(1), (2)].

Note the words actively and systematically. Waiting for complaints is neither. A complaints inbox is a smoke detector that only works once the fire reaches your desk.

The monitoring system runs on a document, the post-market monitoring plan, part of the technical documentation from day one, in Annex IV Part Nine. The Commission must publish guidance on the plan, including a voluntary template, by 2 September 2027 [1, Art 72(3)]. 

Voluntary is the word to hold onto: the final text deliberately removed the harmonised-template empowerment so providers can shape monitoring to their own organisation, which means the five questions below are yours to answer, not a form’s.

Translated into machinery, a monitoring plan answers five questions, and Chapters 11 and 13 already built most of the answers.

What is watched. Performance against the declared accuracy metrics. Drift in inputs and outputs. Override and escalation patterns. Complaints. And segment-level performance wherever the risk file named specific populations.

Where the data comes from. The Chapter 11 logs, the deployer feedback channel, complaints, support tickets, and, for systems sold to others, the contractual data flow this book has now told you to put in the contract three times. This is the fourth. Article 72 says data provided by deployers. Only your contract makes that sentence describe something that actually happens.

Who looks at it, and how often. A named owner, a stated cadence, and thresholds that separate signal from noise. A monitoring plan without thresholds is a dashboard nobody is obliged to believe. And in most organisations, a dashboard nobody is obliged to believe is a dashboard nobody looks at. Beautifully rendered, permanently green, quietly lying.

What happens when a threshold trips. The escalation route, into the risk process, into corrective action, into incident reporting where the facts demand it. The plan connects to the next two sections of this chapter, or it decorates a shelf.

What gets written down. The reviews, the findings, the actions. Monitoring that leaves no record is, for audit purposes, monitoring that did not happen. Chapter 17’s examiner will sample exactly this, with the enthusiasm of someone who has read this paragraph too.

For systems already inside sectoral post-market regimes, medical devices being the obvious case, the Act permits integration into the existing system rather than duplication. One system, both regimes, documented once. Take the win. There are few enough of them.

Serious incidents: Article 73

Then the day something goes wrong. The Act defines a serious incident as an incident or malfunctioning of an AI system that directly or indirectly leads to one of four outcomes: 

  1. The death of a person, or serious harm to their health. 
  2. A serious and irreversible disruption of critical infrastructure. 
  3. An infringement of obligations under Union law intended to protect fundamental rights. 
  4. Serious harm to property or the environment [1, Art 3(49)].

Read limb three again. Slowly. Death and infrastructure are what incident regimes traditionally mean, and most compliance officers can picture them, something catches fire, someone calls an ambulance. 

Limb three has no smoke. It means a recruitment system that discriminated. A credit model that systematically excluded a protected group. A biometric system that misidentified at scale. For the systems most readers of this book run, limb three is the live one. 

If your incident procedure only imagines physical harm, your incident procedure was written for a factory you do not own.

Providers of high-risk systems must report serious incidents to the market surveillance authorities of the Member State where the incident occurred. Now the clock structure. The report is due immediately after the provider establishes a causal link between the system and the incident, or the reasonable likelihood of one, and in any event within fifteen days of awareness [1, Art 73]. 

Two accelerations sit on top. Two days for a widespread infringement or a serious and irreversible disruption of critical infrastructure. Ten days where a person has died. And an incomplete initial report, followed later by a complete one, is expressly permitted.

Deployers sit inside this machinery too. A deployer identifying a serious incident informs the provider immediately, and the duties flow onward from there. Where the provider cannot be reached, the deployer’s own reporting route engages [1, Arts 26, 73]. 

Once more, the statute assumes a pipe between deployer and provider that only your contract actually lays. Fifth mention. The contracts chapter of your vendor negotiation should by now be developing a complex about it.

Three things must be built before the first incident, because the deadlines are far too short to improvise inside, and improvisation under a two-day clock produces prose you will regret.

Name the judge

The clock starts on awareness and causal link, so someone must watch for candidate incidents, and someone must own the causal-link call. 

That call is uncomfortable by design. Reporting concedes a link. Not reporting risks a blown deadline. The person deciding at speed should be named in advance, with counsel’s number written down, not searched for on the day. The permitted incomplete report is the pressure valve. 

When in doubt, file it, and investigate underneath it. A conservative report is recoverable. A missed deadline is a finding with your name on it.

Translate the triggers

Nobody in a credit operations team spots an “infringement of obligations intended to protect fundamental rights” in the wild. They spot a complaint pattern, a segment gap, a journalist asking oddly specific questions. 

The Article 72 thresholds and the Article 73 triggers should therefore be one list, written in the language of the people who will actually see the signal first, and trained into them under Chapter 9.

Pre-build the file 

An incident report needs the system version, the logs for the affected period, the affected population, and the timeline of awareness. Chapter 11 built the reconstruction chain, and the incident procedure is its first paying customer. 

The litigation hold fires on the same trigger, and the procedure should say so explicitly, because deletion schedules do not read the news, and “the logs aged out on Tuesday” is a sentence with a very short future in it.

Corrective action, withdrawal and recall: Article 20

Article 73 told the authority something went wrong. Article 20 is about fixing it.

The rule works like this. If you are a provider and you learn, or have reason to suspect, that a high-risk system you put on the market does not comply with the Act, you must act immediately. 

Four options are open to you: 

  1. Fix the system so it complies. 
  2. Stop supplying it. 
  3. Disable it. 
  4. Pull it back from the deployments where it already runs. 

You choose whichever fits the problem. And you must tell everyone in your supply chain, distributors, deployers, your authorised representative, importers [1, Art 20]. 

If the system presents a risk, the market surveillance authorities hear about it too, what the non-compliance is, and what you are doing about it.

Two of those options carry precise product-law names, and AI discussions blur them constantly, so here is the vocabulary. Withdrawal stops new supply. No new sales, no new customers, no new provisioning, while whatever is already deployed keeps running. Recall reaches what is already out there. Existing deployments are corrected, disabled or returned.

For software, recall is technically the easy part. Roll back the model, force an update, switch off the endpoint. Done by lunchtime. 

The hard part is organisational, because your deployers built their daily operations around the thing you just switched off. Their loan applications still arrive. Their job candidates still apply. So a proper corrective action plan for a decision system includes what the deployers do instead, the manual fallback from Chapter 12, activated across every customer at once, with proper notice.

And then the question everyone forgets. What happens to the people the faulty model already scored? The applicants it wrongly declined? The candidates it wrongly rejected?

“We fixed the model” is an engineering statement. The queue of past decisions is a legal problem, and it does not fix itself. Your plan needs an answer. Re-run the affected cases, notify the affected people, or hold a documented reason why neither is needed.

One last point of context. Everything above describes what you do voluntarily. If you do not, the authorities can order all of it, on their timetable, using Chapter 17’s powers. That is the whole argument for acting early and documenting as you go. 

You stay the author of your own remediation, instead of the recipient of someone else’s. Authors get editorial control. Recipients get deadlines.

Closing the loop: incidents feed the risk system

Now the connection that turns the apparatus into a system rather than a filing habit. Article 9’s risk management process is continuous and iterative, and its inputs expressly include post-market monitoring data [1, Art 9]. 

The loop, in one sentence: monitoring detects, incidents crystallise, corrective action responds, and the risk file updates, so that the next assessment starts from what the field taught you.

The loop closes through four write-backs, and each one is a document an examiner can request.

The risk register. The incident’s scenario gets added, or its likelihood revised, dated, with the pre-incident assessment left visible. The honest record of having been wrong is precisely the evidence that the process works. 

Do not tidy the history. Tidied history is the compliance equivalent of repainting the skid marks.

The technical documentation. The Part Six changelog carries the corrective action, and the affected sections re-version.

The instructions for use. If the field revealed a limitation, deployers hear about it through the document they are legally bound to follow, not through a blog post they are free to miss.

The monitoring plan itself. Thresholds tuned by what they failed to catch. The loop, learning about its own senses.

The anti-pattern has a name in every safety-critical industry: the incident that changed nothing. Report filed, model patched, risk file untouched, and eighteen months later the same failure returns with a different serial number, now decorated with a paper trail proving the organisation had notice. 

Notice plus inaction is the expensive combination, under Chapter 19’s penalty logic and any court’s negligence logic alike. The loop is cheap. Its absence compounds.

What an auditor will ask for

For this chapter, expect the following requests: 

  • The monitoring plan for a sampled system, read against the Commission’s guidance once published, plus the review records showing it runs at its stated cadence.
  • The thresholds, and one example of a threshold tripping, with what followed, because a monitoring system in which no threshold has ever tripped is either watching a perfect system or watching nothing, and the examiner knows which is more common. 
  • The incident procedure, with the named causal-link owner and the operational trigger list. 
  • For any reported incident, the report itself, the timeline from first signal to filing measured against the statutory clocks, the logs under hold, the corrective action, and the notifications. 
  • For any incident assessed and not reported, the assessment itself, because the decision not to report is a record, and its absence reads as a decision not to decide. 
  • Finally, the loop evidence, the risk register entry, the documentation re-version, the instructions-for-use update, showing the incident changed the file and not merely the model.

The regime above assumes things occasionally go wrong, and prices the honesty of saying so. The next chapter covers the pricing when the honesty is missing. Penalties, and the enforcement reality taking shape across the Member States.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 19: Penalties and Enforcement Reality

Every compliance book saves the fines for the end, on the theory that fear works better once you know what the obligations are. This book is no different. You have now read eighteen chapters of duties. Here is what they cost when ignored, who collects, and, since the enforcement record is still young enough to fit on a postcard, an honest reading of who is likely to be fined first.

The three tiers

Article 99 sets three fine ceilings, and the architecture is simple enough to memorise in a lift [1, Art 99].

Tier one: up to 35 million euros or 7 per cent of total worldwide annual turnover, whichever is higher, for breaching the Article 5 prohibitions. The red lines from Chapter 4 carry the red number. For comparison, the GDPR tops out at 4 per cent. The AI Act looked at the regulation that terrified boardrooms for a decade, and raised it.

Tier two: up to 15 million euros or 3 per cent, whichever is higher, for non-compliance with the listed operator obligations. The provider duties of Chapter III, the deployer duties of Article 26, the transparency duties of Article 50, and the rest of the machinery this book spent its middle chapters building [1, Art 99(4)].

Tier three: up to 7.5 million euros or 1 per cent, whichever is higher, for supplying incorrect, incomplete or misleading information to notified bodies or authorities [1, Art 99(5)]. The quiet tier. Do not let its position at the bottom fool you, because it is the one you can breach with a single optimistic email during an inspection. Chapters 16 and 17 warned you about this tier without naming it. Now it has a name and a price.

One note of scope. Article 4, the AI literacy provision from Chapter 9, appears in none of these lists. Whatever the final Omnibus wording did to that obligation, it was never backed by an Article 99 fine band. Any vendor deck telling you otherwise is selling training with menaces.

The SME inversion

Buried in Article 99(6) is the most important sentence in this chapter for smaller readers. For SMEs and start-ups, each fine is capped at whichever of the two figures is lower, not higher [1, Art 99(6)].

Run the arithmetic. A start-up with two million euros of turnover facing a tier-two breach is exposed to 3 per cent of turnover, sixty thousand euros. Not fifteen million. The headline numbers are written for the companies whose logos you know, and the Act then quietly hands smaller operators a different arithmetic altogether. 

The Omnibus reinforced the direction, requiring Member States to weigh the economic viability of SMEs and small mid-caps when setting penalties [24].

Relief, though, is not immunity. Sixty thousand euros is still a bad quarter. The reputational entry is turnover-blind. And nothing in Article 99(6) discounts the corrective action orders of Chapter 18, which can hurt considerably more than the fine.

How the number is actually set

The ceilings are only ceilings. The actual figure is set by the authority weighing the factors in Article 99(7): 

The nature, gravity and duration of the infringement. 

  • Its consequences, and the number of persons affected. Intent or negligence. 
  • Actions taken to mitigate. 
  • Previous fines. 
  • Cooperation with the authority. 
  • How the authority learned of the breach, including whether the operator reported it themselves [1, Art 99(7)].

Read that list as a discount schedule, because that is what it is. Documented compliance effort, early corrective action, cooperation, self-reporting: every one of them moves the number down. Which means the entire evidence architecture of this book, the maps, the records, the closed-loop write-backs, is, among its other functions, mitigation pre-assembled. An organisation with Chapter 17’s file does not just defend better, it also prices better. 

GPAI fines: Brussels collects directly

The model providers of Chapter 15 answer to a different cashier. Under Article 101, the Commission may fine providers of general-purpose AI models up to 15 million euros or 3 per cent of worldwide turnover for breaching Chapter V. These fines are levied centrally, through the AI Office, with the power applicable from 2 August 2026 [1, Art 101] [6].

The Omnibus then extended Brussels’s reach. The AI Office gains market surveillance powers with teeth: 

  • Periodic penalty payments of up to 5 per cent of average daily worldwide turnover, per day of continued non-compliance. 
  • Recovery of its supervisory costs from non-compliant operators. 
  • A five-year limitation period for enforcement [24]. 

The package also opens the door for the Commission to intervene directly in serious cross-border cases [26], a structure familiar from a decade of GDPR politics, where waiting for a national authority became its own genre of complaint.

The national patchwork

Below Brussels, enforcement belongs to whatever each Member State built, and the builds vary widely. Some countries appointed a single AI authority. Others distributed the Act across their existing regulators, with Ireland’s model from Chapter 4 as the instructive extreme. 

Different prohibitions assigned to the financial regulator, the workplace commission, the media regulator and the data protection authority, each carrying its own culture, its own backlog, and its own appetite [12].

The practical consequence for a multi-country operator is that your AI Act exposure has a geography. The same system, deployed identically in two Member States, faces different examiners with different sectoral instincts and different penalty traditions. The evidence map from Chapter 17 does not change. The tempo, and the door the letter arrives through, do. 

Know which authority owns your use case in each market where you matter, and read what that authority has published, because regulators announce their priorities in speeches long before they announce them in decisions.

The enforcement record so far, and how to read it

The honest ledger, as at the time of writing: no published fines under the AI Act, anywhere. But an empty ledger two years into a regime is not unusual, and the GDPR’s own history is the reference case. 

Applicable May 2018. First headline fine in January 2019. The serious money arrives in years three to five. 

Enforcement machinery loads slowly. Authorities staff up, complaints accumulate, investigations mature, and the first decisions get chosen carefully, because first decisions are precedents, and regulators know it.

Which is exactly why the current quiet is strategic information rather than comfort. The complaints being filed now are the decisions of 2027 and 2028. The organisations relaxing now are the respondents of 2027 and 2028. The overlap between those two groups builds quite a pipeline.

The GDPR next door

No AI Act analysis is complete without noticing the older, better-armed regime in the adjacent room. Most consequential AI systems process personal data, which means most AI Act facts are also GDPR facts. And the GDPR enforcement machine is fully warmed up, fully staffed, and holding a decade of case law.

Three practical consequences follow. 

First, dual exposure is the norm, not the exception. A discriminatory screening model is an AI Act problem, a GDPR fairness problem, and possibly an automated decision-making problem under GDPR Article 22, and the fines are separate regimes with separate ceilings. 

Second, the data protection authorities sit inside the AI Act’s structure, designated as enforcers of parts of it in many Member States and formally involved in others. So the examiner who knows your name from a GDPR matter may be the same examiner reading your AI file, with all the institutional memory that implies. 

Third, and this is the near-term reality: while AI Act enforcement loads, GDPR enforcement of AI is already running. The Clearview line of decisions from Chapter 4 was GDPR enforcement of conduct the AI Act now prohibits directly. The authorities did not wait for the new statute, and they will not retire the old one now that it has a sibling.

For the next two years, the most likely regulator at your door about your AI is a data protection authority holding the GDPR in one hand and the AI Act in the other. The file that answers one had better not contradict the file that answers the other.

Who actually gets fined first, and why

Prediction is a risky business to put in print. So what follows is reasoning, not prophecy, and the reasoning is about enforcement economics. First cases get chosen for clarity, visibility and pedagogy.

The prohibition case. Tier one exists to be used, the Commission has a political need to show the red lines are real, and Chapter 4’s grey zones, workplace emotion recognition above all, offer facts that photograph well. 

A prohibition case requires no accuracy benchmarking, no standards debate, no duelling experts. The practice either occurred or it did not. Clean, visible, pedagogical.

The Article 50 case

From August 2026, the transparency duties become the most breached obligations in the Union by sheer volume, and the breaches are verifiable from the outside. An undisclosed chatbot is its own evidence, collectible by an authority with a browser and a free afternoon. Expect sweeps. Expect the first respondents to be chosen for visibility. And expect the fines to be modest and the press releases not to be.

The GPAI case

The AI Office holds direct jurisdiction over a small, famous population with published obligations and, from August 2026, a fining power it will want to demonstrate exists. The first Chapter V matter may well settle into commitments rather than a fine, that is how new regulators often open. The sector should not mistake the opening bid for the ceiling.

The information case (The sleeper)

Somewhere in the first wave of inspections, an operator will hand the authority a declaration, a test report or an answer that its own logs contradict, and tier three will have its debut. This is the fine you earn during enforcement rather than before it, which makes it the purest test of Chapter 17’s disciplines, and the most avoidable entry in the whole schedule.

The deemed provider

Not a separate tier, but the multiplier running through all of them. The deployer who rebranded, substantially modified or repurposed a system, became its provider under Article 25 without noticing, and now faces provider-grade obligations and provider-grade exposure, holding a deployer-grade file. Chapter 2 called Article 25 the place where comfortable role assumptions go to die. Article 99 is where the estate is settled.

The common thread, and the closing argument of this chapter. Every early-case profile above is a failure of the cheap artefacts: 

  • The screen not run. 
  • The disclosure not added. 
  • The analysis not written. 
  • The answer not checked against the logs.

The expensive machinery of this book, the risk systems and the testing regimes, protects you in year five. 

The one-page records protect you in year one. And year one has started.

What an auditor will ask for

For this chapter, the request list is short, because penalties are what happens after the other lists went badly. Expect interest in three things: 

  • Your penalty exposure assessment, which systems, which tiers, which geographies, as evidence that Chapter 8’s governance actually priced the risk. 
  • Your mitigation posture against the Article 99(7) factors, the cooperation procedures, the self-reporting thresholds, the corrective action discipline of Chapter 18. 
  • The consistency of your AI Act file with your GDPR file, because the first cross-regime contradiction an examiner finds is a tier-three conversation waiting to happen.

The Act’s own machinery ends here. The next chapter widens the lens to the regimes it lives beside, because no system in your inventory answers to one statute alone.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 20: Multi-Framework Reality

This book has spent nineteen chapters pretending, politely, that the AI Act is the only law you have. It is time to stop.

No system in your inventory answers to one statute. The credit model answers to the AI Act, the GDPR, consumer credit law, and, if you are a bank, the operational resilience regime. The diagnostic tool answers to the AI Act and the medical devices regulation. The chatbot answers to the AI Act, the GDPR, and, if your platform is large enough, the Digital Services Act. Each regime arrived separately, was drafted by different people, and uses different words for overlapping ideas.

The result, in most organisations, is five compliance programmes running in parallel. Each has its own documents, owners and deadlines. They mostly describe the same systems. They occasionally contradict each other. 

Collectively they produce the problem this chapter exists to solve: nobody can say which document governs what, when. 

This chapter maps the overlaps, then makes the argument the whole book has been building towards. One governance system, many regimes. Or you will drown in your own paperwork.

The GDPR: the sibling you already live with

Start with the regime you already run, because the overlap is largest and the traps are oldest.

The two regulations attach to different things. The GDPR attaches to processing of personal data. The AI Act attaches to systems. But most consequential AI processes personal data, so most systems sit inside both regimes at once, and the practical question is where the files touch. Four contact points do most of the work:

Impact assessments

The GDPR’s data protection impact assessment and the AI Act’s fundamental rights impact assessment under Article 27 examine overlapping harms through overlapping questions, and the Act itself says a deployer’s FRIA may build on an existing DPIA [1, Art 27]. 

Two assessments, one system, one set of facts. The organisations doing this well run them as a single exercise with two outputs. The organisations doing it badly discover, in an inspection, that their DPIA and their FRIA describe two different systems that happen to share a name.

Automated decision-making

GDPR Article 22 restricts solely automated decisions with legal or similarly significant effects. The AI Act’s oversight machinery from Chapter 12 regulates how humans supervise exactly those decisions. 

And here is the trap. Your GDPR position probably claims meaningful human involvement, because that is what takes you outside Article 22. Meanwhile your AI Act file records override rates and time-on-case metrics, because Chapter 12 told it to. If those metrics show rubber-stamping, they do not merely embarrass your oversight design. They demolish your GDPR position, in writing, in a file you are legally required to keep. 

Cross-regime consistency is survival.

Data governance

AI Act Article 10 disciplines training data. The GDPR disciplines the lawful basis for holding it at all. The Omnibus added a bridge worth knowing here, now a pathway in Article 4a, for processing special categories of personal data for bias detection and correction, open to deployers as well as providers [1, Art 4a]. 

This resolves the old paradox that measuring discrimination required exactly the sensitive data the GDPR discouraged you from keeping. 

Chapter 7 covered the conditions, and they are strict. The bridge exists. It is narrow, and it is inspected.

Enforcement

Chapter 19 said it, and it bears one repetition. For the next two years, the likeliest regulator at your door about AI is a data protection authority holding both statutes. One system, one authority, two files. They will be read together. Write them together.

The DSA: the platform layer

The Digital Services Act regulates intermediary services, and its AI overlap concentrates at the top of its pyramid. Very large platforms and search engines owe systemic risk assessments covering their recommender systems and content moderation, transparency about recommendation parameters, and researcher access to data.

If you are not a platform, the DSA touches you lightly. But two of its ideas travel well. 

Recommender transparency is Article 50’s cousin, because both regimes are converging on the same principle, that people should understand what the machine is doing to their information diet. 

The DSA’s systemic risk assessment is the FRIA’s older sibling, working at population scale. 

Organisations inside both regimes should notice something. Three different laws now ask them to assess roughly the same algorithmic harms. That is either three reports, or one honest analysis wearing three cover pages. Choose the second.

Sector rules: where the Act was designed to nest

The AI Act was built knowing it would land on top of sectoral regimes, and Chapter 16 showed the mechanism. Annex I systems take their conformity assessment through the sectoral route. One examination, both rulebooks.

Financial services

The dense corner. A bank’s credit model sits under the AI Act as Annex III high-risk, under the GDPR, under consumer credit rules on creditworthiness assessment, and inside an institution regulated for model risk and operational resilience, where DORA governs the ICT resilience of the estate the model runs on. 

The good news is that financial institutions already run model governance, validation functions and risk committees, so Chapter 6’s risk system has somewhere natural to live. The bad news is that the existing machinery speaks a different vocabulary. 

Model validation is not conformity assessment, and mapping the two without doing everything twice is precisely the “which document, when” problem below. The Act helps once, and expressly. 

For financial institutions, certain AI Act governance obligations may be discharged through compliance with financial services governance rules [1, Arts 17, 26]. That is an integration invitation, and the sector should accept it in writing.

Medical devices

The cleanest nesting in the whole landscape. An AI diagnostic is a medical device under the MDR and an Annex I high-risk system under the AI Act. The notified body examination absorbs the AI requirements. 

The technical documentation is one file serving both regimes. The post-market surveillance systems merge, under the express permission Chapter 18 noted. 

Device manufacturers have run this kind of regime for decades, and they mostly need translation rather than transformation. Chapter 10’s living documentation is their technical file, with new sections.

Machinery

The moving target. The Omnibus took the Machinery Regulation out of Annex I Section A, dissolving the dual-assessment model [4]. AI-specific requirements will arrive instead through delegated acts inside the machinery regime, by August 2028. 

For machinery makers this is genuine simplification, with a planning cost attached. The requirements are coming, their content is not yet written, and the trigger list from Chapter 5 should be watching for the delegated acts by name.

The “which document, when” problem

Now the disease itself. Take one system, a credit decisioning model at a mid-sized lender, and list what the regimes demand. A DPIA. A FRIA. An Article 6 classification record. Technical documentation, or its deployer-side shadow. The instructions-for-use mapping. An oversight design. A logging specification. A monitoring plan. A model validation file. A DORA register entry. A creditworthiness assessment procedure. A records-of-processing entry.

Twelve documents. Five regimes. One model.

Now ask the questions an examiner asks. Which document is the authoritative statement of the system’s purpose? If the DPIA says decision support and the classification record says material influence on outcomes, which one is true, and who noticed? When the model retrains, which documents update, in what order, and who knows the order exists?

Most organisations cannot answer. The documents were born in different departments, in different years, and each one is somebody’s finished project rather than everybody’s living record. Understand the failure mode precisely. It is not missing paperwork. It is abundant, contradictory paperwork, which is worse, because every contradiction is discoverable, and Chapter 19’s tier three is fed by exactly this.

The cure is boring and total: 

A document map per system, the multi-regime extension of Chapter 17’s evidence map. For each system, every required artefact across every applicable regime, with four annotations:

  • Which regime demands it. 
  • Which document is authoritative for each shared fact, purpose lives in the classification record, and everything else cites it. 
  • Which events update it. 
  • Who owns it. 

One page per system. The exercise of writing it is where the contradictions surface, which is, of course, the point of writing it.

One governance system instead of five

The map reveals the deeper truth. The five programmes are one programme wearing five costumes.

Every regime in this chapter reduces to the same skeleton. Know what you run: an inventory. Know what applies to it: scoping and classification. Know what that demands: obligations. Prove you do it: evidence, owned and dated. Notice when anything changes: triggers, for the system, the data, the vendor, and the law itself. 

The GDPR calls this a register of processing and DPIAs. The AI Act calls it classification and technical documentation. DORA calls it a register of information. The MDR calls it a technical file. Different nouns. Same machine.

So build the machine once. One inventory, where every system appears exactly once, carrying its roles and classifications under every applicable regime. One evidence layer, where each artefact exists once, is owned once, and is cited by every regime that needs it, so the DPIA and the FRIA share their facts instead of quietly diverging from them. One trigger discipline, where a model retrain, a vendor swap or a legislative amendment fires the updates across every affected regime in a defined order, because Chapter 3 taught you that the law itself is one of the things that changes. 

And one reporting line into Chapter 8’s governance, because a board asked five separate compliance questions will eventually notice they are the same question, and wonder why they are paying to answer it five times.

What I built, because this chapter kept happening

A disclosure, and then a short pitch, clearly labelled as one. I have spent years inside the problem this chapter describes, as a regulatory lawyer and as a founder, and the tooling I wanted did not exist.

Compliance software mostly falls into two camps. Questionnaire products ask your team what it believes and file the beliefs. Document platforms store the five costumes in five folders, contradictions included.

So I am building the third thing. Its design follows this book because both follow the same conviction: compliance is an evidentiary state, assessed at the level of a specific system, across every regime that touches it.

In my platform, a system exists once. Its obligations are scattered across regimes — AI Act, GDPR, ISO 42001 and onwards — derived from recorded facts about the system, not from a self-declared risk label or what the organisation calls itself.

Every obligation joins to its artefact, its owner and its location. Chapter 17’s evidence map as a living object rather than a spreadsheet that dies quarterly.

When the law moves, the Omnibus being the recent demonstration, the change is designed to propagate as a framework event: every affected system re-scoped, every affected obligation re-fired, every stale document flagged, in order, with a trail. The intention is that the mock audit of Chapter 17 becomes a button rather than a fortnight.

That is the pitch, and it is the whole pitch. I should be plain about the stage: my platform is in build, working with a small number of design partners, and it is not yet generally available. 

The method in this book does not wait for it. The method works on paper, with discipline and a calendar. Every artefact described in these chapters can be built in ordinary documents, and Chapter 21 shows the sequence.

Software earns its place when the inventory grows past what discipline alone can hold consistent. If that is where you are, europeancompliancesuite.com is where the templates live.

End of advertisement. The book now returns to its regularly scheduled impartiality.

What an auditor will ask for

For this chapter, expect the following requests, and note that they can arrive from any of five directions, or more: 

  • The document map for a sampled system, across all applicable regimes. 
  • The consistency of shared facts, purpose, population, oversight reality, as stated in the DPIA, the FRIA, the classification record and the GDPR file, read side by side, because they will be. 
  • The integration evidence wherever regimes permit merging, the combined assessment, the merged monitoring system, the single technical file, each with the permitting clause cited. The update trail for the last material change, which documents moved, in which order, under whose hand. 
  • For financial institutions, the mapping between model governance vocabulary and AI Act vocabulary, because the examiner who speaks both dialects has already started work, and the first thing they check is whether you do.

One chapter of machinery remains: the sequence. Chapter 21 turns the whole book into a ninety-day plan.

This chapter’s auditor checklist is available as an editable template, free, at europeancompliancesuite.com/eu-ai-act-toolkit 

Part V — Making It Operational

Four parts of law, machinery and evidence. What remains is the question every reader has been storing up since Chapter 1. Fine, but what do I do on Monday?

Part V answers it. Chapter 21 turns the book into a sequenced plan, ninety days, three company profiles, ordered by what actually blocks what. Chapter 22 closes with the failure catalogue, the mistakes that recur so reliably they deserve names, and the corrective for each one. Nothing new is introduced here. Everything is arranged for immediate use. If Parts I to IV were the anatomy lesson, Part V is the part where you operate.

Chapter 21: The 90-Day Compliance Sprint

No organisation implements this book in ninety days. That is not the claim, so let us retire it in the first paragraph. A risk management system with review history, a testing regime keyed to versions, an oversight design with behavioural metrics, these take quarters to build and longer to season. And Chapter 3 explained why the seasoning itself is the evidence.

What ninety days can do is different, and more valuable than it sounds. Establish the foundations everything else depends on. Kill the risks that cannot wait. And leave you with a calendar instead of a cloud of anxiety. The sprint is not the programme. It is the part of the programme that unblocks the rest of the programme.

The sequencing principle is dependency, not importance. Ask “what matters most” and everything matters most. That is exactly how organisations end up eight workstreams wide and zero workstreams deep, with a steering committee and no inventory. 

Ask instead “what does everything else depend on”, and the order writes itself. 

You cannot classify systems you have not inventoried. You cannot map obligations for systems you have not classified. You cannot collect evidence for obligations you have not mapped. And you cannot do any of it twice as well under enforcement pressure as you can now, calmly, at a fraction of the price.

One rule governs all three profiles before a single task starts. Name the owner. Not a committee, a person. Ninety days is short enough that shared ownership means no ownership, and every task below assumes somebody wakes up responsible for it. 

If your organisation cannot name that person by day three, that is not a planning obstacle. That is finding one, and it goes to whoever runs Chapter 8’s governance question first.

The chapter runs three profiles, because the same ninety days buys different things depending on who you are. A deployer-only organisation is building visibility and vendor discipline. A provider SME is building the file that self-certification will one day sign. A scale-up crossing into provider territory is doing both at once while the ground moves underneath it, which is why its profile comes last and reads fastest. 

Find yourself, and ignore the other two sprints without guilt. That is what profiles are for.

Profile one: the deployer-only organisation

You build nothing, sell nothing AI, and buy everything. The HR platform with the screening module. The chatbot from a SaaS vendor. The credit or fraud scores arriving through an API. 

Most organisations are you. Your sprint is about discovering what you actually run, establishing what it owes, and converting a drawer of vendor contracts into a set of answers. Nothing here waits for December 2027. Most of it was due, in spirit or in law, some time ago.

Days 1 to 15: find everything

The inventory, Chapter 8’s master record, comes first because every later step consumes it. Survey the business units, but do not rely on the survey, because people do not report AI they do not know is AI, which after Chapter 2 you know is most of it. 

So triangulate. The procurement ledger for software spend. The IT asset register. The DPIA log. Marketing’s tool stack. And one pointed question to every department head: what makes recommendations, scores, rankings or decisions around here? 

Expect the count to come in at two to three times the first guess, and expect at least one surprise that changes the rest of the sprint. 

Record per entry what it does, who runs it, which vendor supplies it, and which population it touches. Resist depth. This fortnight is a census, not an audit.

Days 16 to 35: sort everything

Classification, at sprint speed. First pass, the definitional screen from Chapter 2, using a one-page version of the classification policy. Rule-based tools out, learned systems in, doubtfuls flagged rather than debated. 

Second pass, the prohibition screen from Chapter 4 across everything that survived, with the emotion recognition question asked explicitly of every workplace analytics and productivity tool, because that is where the bodies are buried. 

Third pass, Annex III mapping. Which systems touch recruitment, worker management, credit, essential services. Full reasoned records can follow in the months after. 

The sprint deliverable is the sorted list with one-paragraph rationales and a flagged-for-analysis queue, dated, because Chapter 5 taught you that a dated wrong-ish record beats an undated silence.

Days 36 to 60: interrogate the vendors

Now the deployer’s core discipline gets its workout. For every system in an Annex III area and every customer-facing AI, pull the contract and the documentation, and answer four questions. 

  1. Do we hold the instructions for use, current version? 
  2. What does the contract say about logs, custody, export, retention, Chapter 11’s question? 
  3. Who implements the Article 50 disclosures, us or them, in writing, Chapter 14’s question? 
  4. What change notice do we get when the model under our process silently becomes a different model, Chapter 15’s question?

The answers will be mostly gaps. Good. The sprint deliverable is the gap list converted into vendor letters, sent, with responses tracked. 

You are not renegotiating every contract in ninety days. You are creating the paper trail showing you asked, which is worth more than it sounds, and costs a stamp.

Days 61 to 80: fix what is visibly broken

Two categories jump the queue, because they are already in force. Article 50 first. Every chatbot and customer-facing AI gets its disclosure, and every published synthetic media workflow gets its labelling rule, per Chapter 14, because August 2026 does not care about your sprint. Then oversight-in-fact for the consequential systems. For each Annex III-area system, name the human overseers, check they exist, check they know, and give them the short system-specific briefing Chapter 9 described. 

Not the full literacy programme, the triage version. The people standing next to the highest-risk machines learn what the machines cannot do, this month.

Days 81 to 90: institutionalise

The closing fortnight converts sprint output into standing machinery: 

  • The evidence map, started, obligations to artefacts to owners to locations for the top ten systems, Chapter 17’s format, with the gaps marked honestly. 
  • The compliance calendar, loaded, Article 50 done, FRIA obligations diarised where they will bite, vendor response deadlines, the December 2027 runway with its lead times, Chapter 3’s structure. 
  • The triggers, assigned, so that new tool procurement flows through the classification screen from now on, which is the single control that stops the inventory rotting. And the governance slot, booked. 

Thirty minutes, quarterly, with whoever owns Chapter 8’s question, reviewing the map, the calendar and the gap list.

Day ninety’s deliverable is not a compliance certificate in a cute frame. It is a system that knows what it runs, what it owes, what is missing, and when it will be fixed. That alone puts you, without exaggeration, ahead of most of your market.

Profile two: the provider SME

You build an AI system and sell it, and somewhere in Annex III your use case is waiting for you. Your sprint differs in kind from the deployer’s. They are discovering what they bought. You are starting the file that, one day, someone in your company signs a declaration of conformity on top of. Chapter 16 explained what that signature means. The sprint’s job is to make it eventually honest.

The good news first, since providers get little of it. You have one system, or a handful, not a sprawling estate. The census the deployer spends a fortnight on takes you an afternoon. 

So your ninety days go into depth, and into the one asset that appreciates fastest, development evidence captured while development is actually happening.

Days 1 to 10: settle identity

Three documents, none of them long, all of them load-bearing. 

The intended purpose statement, what the system is for, on which population, in which context, written once and made canonical, because Chapter 10 warned you that a purpose stated four slightly different ways across four documents is the first inconsistency an examiner finds. 

The classification record, the Article 6 analysis in full, the Annex III entry, the Article 6(3) filter honestly considered, the profiling override addressed, dated and signed, per Chapter 5. 

The role map: provider of what, exactly, including the GPAI question, whether the model underneath is yours, licensed, or fine-tuned into ambiguity, with Chapter 15’s compute analysis done now, before anyone trains anything else.

Days 11 to 30: start the repository

Chapter 10’s single source, opened this month, not in 2027. The Annex IV structure, nine parts, created as a skeleton. The parts development is currently generating, design choices, data provenance, test results, captured as they occur, keyed to versions, with owners named. 

Do not write the parts you cannot yet write. A skeleton with honest gaps beats a narrative written backwards later, and Chapter 3 already told you why the timestamps themselves are the asset. 

Alongside it, the logging specification, Chapter 11’s one page, drafted now while logging is still cheap to change. Retrofitting traceability into a shipped architecture is the expensive version of every decision you could make this fortnight.

Days 31 to 55: stand up the risk system, small

Article 9 says continuous and iterative. It does not say enormous. A risk register for the one system, populated in a structured session with engineering and whoever knows the deployment context. Failure modes, affected persons, and the oversight and design measures answering each. 

The first review goes in the diary. The output feeds three places at once, the repository’s Part Five, the oversight design, and the accuracy conversation, which happens to be the sprint’s next stop. That triple use is why the risk session earns its week.

Days 56 to 75: face the metrics

Chapter 13’s declaration discipline, started early, because it cannot be rushed later. Choose the metrics that fit the harm, with the error directions separated. Establish test data that resembles deployment, per-group where the risk register said so. Run the first structured evaluation and file the report, dated, versioned, failures included. 

You are not declaring conformity in ninety days. You are discovering, while it is still a private discovery, whether the system performs the way the sales deck says it does. Some sprints end differently than they began at this step. Better now, in your own conference room, than in Chapter 18’s incident procedure.

Days 76 to 90: point outward

The instructions for use, first draft, generated from the repository, with Chapter 12’s four jobs in mind and the limitations written as operational sentences: 

The Article 50 duties, live in August 2026, implemented, the interaction disclosure if your product talks to people, the marking plan if it generates content, per Chapter 14’s timeline, which for a product launching after August 2026 means compliance at launch, not December. 

The conformity roadmap: Internal control under Annex VI mapped as a project, the standards watch assigned, the December 2027 runway drawn backwards, with the notified-body question answered if biometrics are anywhere near you.

Day ninety’s deliverable: a repository that fills itself as you build, a risk loop that runs, honest first numbers, and a dated trail proving all of it started early. 

That trail is worth money in every due diligence you will ever sit through. For an SME provider, that is the sprint’s quiet second purpose.

Profile three: the scale-up becoming a provider

You started as a deployer, or as something unclassifiable, and success is converting you. The white-label deal is signed. The fine-tuning went deeper than anyone documented. 

The “internal tool” has three external customers. Chapter 2 named the mechanism, Article 25, and you are its case study in progress. 

Your sprint is the shortest to describe and the hardest to run, because it is both previous sprints compressed, plus one question neither of them faces.

Days 1 to 20: establish what you have become. 

The Article 25 analysis, in writing, before anything else, because every subsequent obligation depends on the answer. 

For each product line: whose name is on it, what was modified, how much, against which documented perimeter, with what compute, and what the intended purpose has drifted into. Involve counsel, there is no way around it. This is Chapter 2’s judgement call with your company’s entire obligation set riding on it. 

The possible outcomes, still a deployer, a provider of a system, a provider of a model, or several at once across different products, are not equally expensive. 

Knowing which you are is the sprint. Everything after this is a consequence.

Days 21 to 90: run the profile the analysis assigned you, with three scale-up-specific overlays.

First, the paper catches up with the build. Your documentation debt is the gap between what was developed and what was recorded, and the repository you open in week four starts with an honest backfill, labelled as reconstruction wherever it is one. Chapter 10 taught you that a reconstruction pretending to be contemporaneous is worse than the gap it hides.

Second, the contracts catch up with the roles. The white-label agreement allocating provider duties. The upstream model terms checked for version pinning and training-use clauses. The customer contracts checked for what you have already promised, in warranty language Chapter 16 would recognise on sight.

Third, velocity gets a gate. You ship weekly, and from now on the release process carries the classification trigger, the documentation update, and the re-test. Chapter 10’s gate, installed while the habit is still cheap. A scale-up that installs the gate at ten systems thanks itself at fifty.

Day ninety, for you, is a diagnosis and a debt schedule. What you are, what that owes, what is missing, and the dated plan for paying it down. Investors will read it in every round from now on. Write it as if they are already reading.

Closing the sprint

Three sprints, one shape. Days one to ninety buy the same four things in every profile. An inventory that is true. Classifications that are dated. The visibly-broken things, fixed. And a calendar with owners, where the anxiety used to be.

What they do not buy is completion, and the honest close of this chapter is the reminder that day ninety-one exists. 

The mock audit of Chapter 17 belongs around day one hundred and twenty, when the sprint’s artefacts have settled and their gaps are worth finding. The second pass of everything, the full classification records, the literacy matrix, the seasoned risk reviews, fills the quarters the deferral handed you. 

Chapter 3 made the argument once, and it closes the book’s practical case now. The organisations that treat December 2027 as a start date will spend it reconstructing what you, ninety days from reading this, will simply be maintaining.

What remains is the failure catalogue. The mistakes that survive every warning, collected in one place, so that at least they cannot claim to be surprises. Chapter 22, and then you are done, and so is the book.

The consolidated evidence checklist for the whole book is available as an editable toolkit at europeancompliancesuite.com/eu-ai-act-toolkit 

Chapter 22: Common Failure Modes

Every compliance discipline eventually produces its own pathology textbook. The mistakes made so often, by such different organisations, that they stop being accidents and become patterns. AI governance is young, but the patterns are already stable, because none of them is really about AI. They are about wishful thinking, borrowed confidence and theatre, which are all considerably older than software.

This chapter collects the expensive ones. Each comes with its signature, the way it looks from inside while it is happening, and its correction, which by now points to a chapter you have already read. Nothing here is new. That is the point. By the time these mistakes get made, they have all been warned about, usually in the organisation’s own documents.

Classification by wishful thinking

The signature. The classification analysis starts from the answer. Somebody senior has decided the product must not be high-risk, because high-risk means cost, delay and awkward sales conversations, and the analysis is then written backwards from that conclusion. 

The tell is in the language. The system “merely assists”. It “only recommends”. It “supports but never decides”. Phrases doing suspiciously heavy lifting, usually in documents where the sales material simultaneously boasts that the same system “automates decisions end to end”. Marketing and compliance describing one product as two opposite products is the pattern’s fingerprint, and Chapter 17 teaches you who eventually reads them side by side.

Why is it expensive? A wrong classification is not one mistake. It is the seed of a hundred. Every obligation not mapped, every artefact not built, every test not run flows from it. And the discovery usually arrives through Chapter 18’s incident machinery, at which point the file must be built retrospectively, under enforcement pressure, with timestamps that confess everything.

Here’s how to remediate, and the key is in Chapter 5, whole. Classification as a process, with a policy, standing positions, and a named owner whose judgement is insulated from the revenue forecast. 

The nearest-miss discipline, writing down which Annex III entry came closest and why the system falls outside it, which is precisely the sentence wishful thinking cannot write honestly. 

Here’s one piece of cheap insurance. When the call is genuinely close, price both branches before choosing. Sometimes the high-risk programme costs less than the argument, and nobody ever prices it.

Vendor-claim reliance

The signature. The compliance file for a procured system consists of the vendor’s statement that the system is compliant. Sometimes laminated. The organisation has confused buying a product with buying an answer, and the confusion is comfortable for everyone involved. The vendor enjoys asserting. The buyer enjoys believing. Procurement enjoys closing.

Why is it expensive? Chapter 1 said it in week one. Deployer obligations are direct. Your oversight duty, your literacy duty, your disclosure duties, your log custody, none of them transfers with an invoice. 

The vendor’s certificate answers the vendor’s obligations. When the authority writes to you, it writes about yours, and “we relied on the vendor” is the opening line of a negligence finding, not a defence. 

The Article 25 variant is still crueller. The organisation that customised, rebranded or repurposed the vendor’s system has been the provider for months, relying all the while on compliance claims made by a company that is no longer, legally speaking, the provider of anything it runs.

Here’s the remediation: the due diligence method from Chapters 12 and 15. Read what is missing, not what is there. The four vendor questions from Chapter 21’s deployer sprint, asked in writing, with the gap list converted into contract terms, instructions for use current, log custody and export, disclosure allocation, change notice. 

One internal rule that survives every procurement cycle. A vendor claim is an input to your record, never a substitute for it. One page of your own analysis on top of their assertion converts borrowed confidence into owned evidence. One page is all it takes.

Documentation theatre

The compliance binder is magnificent. Two hundred pages, consistent fonts, a table of contents with heroic granularity, produced in a six-week effort everyone remembers with pride. It describes the system as it stood on the day the binder was finished. 

Which is to say, it describes a system that no longer exists, because the model has retrained twice since, and nobody told the binder. 

The theatrical variant has a second act. Templates purchased, filled with the vendor’s example text, company name substituted throughout, and the word “ACME” surviving quietly in paragraph four.

Why is it expensive? Dead documentation fails twice. It fails as compliance, because Article 11 says kept up to date, and an examiner samples for exactly the drift the binder embodies, the declared accuracy of a superseded model, the named overseer who left in January. 

And it also fails as protection, because in an incident the beautiful binder becomes the claimant’s exhibit. The organisation knew precisely what good governance looked like, wrote it down, and did not do it. Half-decent living records beat immaculate dead ones in every forum that matters. The immaculate dead ones just photograph better in judgments.

Here’s how to remediate: start with chapter 10’s five design choices, of which two do most of the work. Single source with generated exhibits, so nothing is written twice and nothing diverges silently. Then, the release gate, so the system cannot change without its paper changing. 

Add the divergence review, twenty minutes on a cadence, which exists because theatre is not a moral failing but a maintenance failure which yields to calendars.

Literacy as a webinar link nobody opened

The training obligation was discharged by an all-staff email containing a link, a forty-minute recording, and a compliance dashboard showing 31 per cent completion, unvisited since. The oversight personnel for the credit model received the same generic module as the canteen team, on the theory that fairness means everyone learning equally little. Attestation consists of the LMS marking “viewed”, a status the software awards for opening the tab.

Why is it expensive? Chapter 9 walked through what survives of Article 4, and the surviving exposure was never really the literacy article anyway. It is Article 26(2), which requires competent, trained and supported oversight of humans for high-risk systems. 

It is the incident inquiry, where the first question is always what the person at the controls had been taught about the machine. A training record showing that the overseer of the screening model completed “AI Awareness Fundamentals”, generic, unversioned, eighteen months stale, answers that question in the claimant’s favour. 

The webinar link is the especially visible symptom of an organisation that trained nobody for the systems that could actually hurt someone.

The correction is simple: Chapter 9’s matrix. Roles mapped to tiers, system-specific content for the people standing beside the consequential systems, attestations that name the system and the version, and the effectiveness layer, escalations made, questions asked, behaviour changed, which is what turns training from a ceremony into a control. 

Sized honestly, the whole apparatus for a mid-sized deployer is days of work. The webinar link was never cheaper, just earlier.

The remaining five, briefly

Five more patterns recur often enough to deserve names, and each one got its full treatment earlier in the book.

The stood-down programme

The deferral read as a holiday, the team dispersed, and eighteen months of unbuildable evidence history quietly forfeited. Chapter 3 made the argument, and the corrective is its sequencing logic. Foundations now, the standard-dependent layers pencilled into the calendar.

The role assumed, never determined

“We are a deployer” held as corporate identity rather than reached as a per-system conclusion, and discovered to be false on the one white-labelled product. Chapters 1 and 2 hold the answer. The corrective is the role determination row in every inventory entry, revisited on Chapter 5’s triggers.

Oversight as furniture

The human in the loop with a 0.3 per cent override rate and three seconds per case, decorating an automated system for regulatory effect. Chapter 12 covers it, and the corrective is the behavioural layer. Measured, reported, and acted on at least once, visibly.

The contract that says nothing

Log custody, disclosure allocation, change notice, incident channels. All of them statutory assumptions, all of them absent from the signed document, all discovered during the first incident, which is the most expensive moment for contract drafting yet invented. 

Chapters 11, 14, 15 and 18 each said it plainly. The corrective is the deployer sprint’s vendor fortnight, run before the incident rather than during it.

The incident that changed nothing

Report filed, model patched, risk file untouched, and the same failure back within two years, now carrying a paper trail that proves notice. Chapter 18 is the guide. The corrective is the four write-backs, plus the internal metric that flags any incident with no linked file changes. That number, incidentally, is the most honest one a governance dashboard can show.

Coda: the file you can hand over

One test gathers every chapter of this book, and it is the test this chapter’s failures all flunk, each in its own way. Imagine handing your complete AI file, the inventory, the classifications, the evidence map, the logs, the training records, the incident history, to a stranger with legal powers and a sceptical disposition. Then leaving the room. Not defending it. Not narrating it. Not explaining what the missing pieces would have said. The file speaks alone.

Wishful classification fails that test because the file argues with itself. Vendor reliance fails because the file is somebody else’s. Theatre fails because the file describes a system that is not the one running. The webinar link fails because the file shows nobody was taught to drive.

Compliance you can prove was the promise of the introduction, and it turns out to be a discipline rather than a destination. Dated records. Honest negatives. Named owners. Closed loops. Maintained on triggers, by people who know they are named. 

None of it is glamorous. All of it is buildable by the organisation you already have, starting, as Chapter 21 put it, on Monday.

The law will keep moving. The Omnibus was not the last word, and this book is a snapshot of a moving target, dated accordingly. 

But the method does not move. Know what you run. Know what it owes. Prove you do it. Notice when anything changes. Everything else lives in the annexes.

Thank you for reading. Now go and open the drawer where the vendor contracts live.

You already know what is not in there.

About the Author

Yuliia Habriiel is a Ukrainian-Canadian regulatory lawyer specialising in European technology regulation, with particular depth in the EU AI Act, which she has worked with since before it had a final article numbering. She took her law degree in Ukraine before her work carried her to Brussels, and eventually into the Act itself, contributing to legislative drafting work during the Act’s development. 

That experience, watching the text argued into shape, shaped the conviction this book is built on: that the distance between what a law says and what an organisation can prove is where compliance actually lives.

Since then she has worked on the implementation side of the same statute. She founded an AI compliance automation platform building audit-ready evidence systems on deterministic legal logic, holds two UK patent applications in the field of automated compliance evaluation, and has advised organisations from startups to established institutions on classification, governance and evidence design. She is the founder of European Compliance Suite (europeancompliancesuite.com), a compliance and advisory practice, and the multi-regime AI governance platform described, with appropriate disclosure, in Chapter 20.

Away from regulation, she writes fiction. Her forthcoming horror short story collection draws on the war in her native Ukraine, on the theory that some things are best understood through the genre built for them. She is a mum of two, and works in English, German, and Ukrainian, though not usually all three before her first morning coffee.

This is her first non-fiction book. 

Appendix: The Consolidated Evidence Checklist

Every chapter of this book closed with the same section: what an auditor will ask for. This appendix collects all of it in one place.

Two notes before you use it. First, where an artefact appears in several chapters, it is listed once, with every chapter that relies on it. Those repeats are not editorial untidiness. They are how examination works: an inspector triangulates by asking for the same document through different doors, and a file that answers consistently through every door is the file this book exists to help you build. 

Second, not every item applies to every reader. A deployer with no generative systems can strike the marking rows. A provider with no biometric systems can strike the two-person verification. Strike honestly, in writing, and the strikes themselves become your scoping record.

The checklist follows the book’s structure: foundations, governance, technical safeguards, then examination and enforcement.

Foundations: inventory, classification, roles

  1. AI system inventory, versioned, covering every system built, bought or embedded, with completeness controls routing procurement, IT onboarding and expense approval into it (Chapters 1, 8, 17)
  2. Classification policy stating the standing positions: rule-based systems, statistical models, optimisation software, embedded components, filter conditions, vendor claims (Chapters 2, 5)
  3. Classification record per system, positive and negative, with nearest-miss reasoning, reconciled exactly to the inventory (Chapters 1, 5)
  4. Definitional analysis per system against the elements of Article 3(1) (Chapter 2)
  5. Role determination per system, with supporting facts (Chapters 1, 8)
  6. Article 25 analysis for every configured, retrained, fine-tuned, rebranded or repurposed system (Chapters 1, 2)
  7. Contractual allocation of provider duties in white-label arrangements (Chapter 2)
  8. Safety component analysis under the post-Omnibus definition, with failure-mode reasoning, for Annex I products (Chapter 5)
  9. Article 6(4) documented assessment per filter exemption: condition relied on, profiling check addressed (Chapter 5)
  10. Simplified registration entry for filter-exempted systems, once the duty bites (Chapter 5)
  11. Version chains for reclassified systems, plus the trigger log showing reclassification fired on events, including regulatory change (Chapter 5)
  12. Handover record from each high-risk conclusion into the compliance programme (Chapter 5)

Calendar and programme

  1. Compliance calendar: obligations mapped to systems, owners, lead-time breakdowns, evidence columns, with its version history (Chapter 3)
  2. Omnibus impact assessment: which dates moved, which held, who signed off the revised plan (Chapter 3)
  3. Pause decision record and restart plan, for any programme stood down in 2026 (Chapter 3)

Prohibitions

  1. Prohibition screen record per system, dated, owned, versioned (Chapter 4)
  2. Analysis memoranda for flagged systems, applying the elements of the relevant prohibition to facts (Chapter 4)
  3. Emotion recognition analysis with deployment constraints, for workplace or education deployments with affect-related functionality (Chapters 4, 14)
  4. Data provenance and consequence map for scoring systems (Chapter 4)
  5. Misuse safeguards assessment for generative systems against the December 2026 prohibitions, plus evidence the safeguards operate (Chapter 4)
  6. Record showing conditional-clear constraints reached the teams able to breach them (Chapter 4)

Risk management

  1. Risk management plan: owner, date, cadence, event triggers (Chapter 6)
  2. Risk register per system, covering intended use and reasonably foreseeable misuse (Chapter 6)
  3. Considered-and-rejected misuse list (Chapter 6)
  4. Linkage records showing post-market data entering the review cycle (Chapters 6, 18)
  5. Measures register, each control traced to an identified risk, in hierarchy order: design first, controls second, warnings last (Chapter 6)
  6. Residual risk acceptance record: named decision-maker, date, criteria, reasoning (Chapters 6, 8)
  7. Vulnerable groups analysis wherever Article 9(9) is engaged (Chapter 6)
  8. Test plan with prior-defined metrics, dated before the test records that report against it (Chapters 6, 13)
  9. Version chain across the whole risk file (Chapter 6)

Data governance

  1. Data governance procedure covering the full Article 10(2) practice list (Chapter 7)
  2. Datasheet per data set per version, with the training, validation and testing split rationale (Chapter 7)
  3. Representativeness analysis tied to named deployment settings (Chapter 7)
  4. Known-error register with tolerance decisions, and the completeness gap register (Chapter 7)
  5. Bias examination records: methods, findings, mitigations, feedback-loop analysis (Chapter 7)
  6. Article 4a file where the pathway was used: necessity analysis, alternatives assessment, safeguard implementation, access log, deletion record, matching GDPR records entry (Chapter 7)

Governance and people

  1. AI policy, dated, board-approved (Chapter 8)
  2. Appointment records: accountable owner, working group terms of reference, system-level owners (Chapter 8)
  3. Board packs and minutes showing the reporting rhythm ran (Chapters 8, 12)
  4. Literacy policy and role-to-tier matrix (Chapters 8, 9)
  5. Training log reconciled to the current staff list and the inventory, module versions included (Chapter 9)
  6. Versioned training content, available for spot review (Chapter 9)
  7. Signed attestations naming the system and version covered (Chapter 9)
  8. Outsourced-operator contractual training clause, plus the evidence returned under it (Chapter 9)
  9. Pre-Omnibus training records, from the period the ensure formulation applied (Chapter 9)
  10. Training records of the persons involved in any incident (Chapter 9)
  11. Worker and affected-person notifications for workplace deployments of high-risk systems (Chapter 8)
  12. Fundamental rights impact assessment with its authority notification, where Article 27 applies (Chapter 8)

Technical documentation

  1. Technical documentation package per high-risk system, current, structured to Annex IV (Chapter 10)
  2. Historical package versions, retrievable as at any specified past date, mapped to system versions (Chapter 10)
  3. Section ownership record, with edit histories showing owners maintain their sections (Chapter 10)
  4. Release-gate evidence: the trigger list in the release process, sampled against the Part Six changelog (Chapters 10, 13)
  5. Intended-purpose reconciliation: Part One against the classification record, the instructions for use and the marketing (Chapters 10, 17)
  6. Simplified documentation form, completed, plus underlying evidence, for SMEs and small mid-caps using it (Chapter 10)

Logging

  1. Logging specification per system, one page, co-owned by system owner and compliance (Chapter 11)
  2. Cross-check of the specification against the technical documentation’s logging description (Chapter 11)
  3. Sampled transaction reconstructed end to end: input, version, output, human action, timestamps (Chapter 11)
  4. Retention schedule with per-category reasoning, plus evidence deletion actually runs (Chapter 11)
  5. Access model, and the access logs on the logs themselves (Chapter 11)
  6. Contractual log custody allocation for procured systems (Chapter 11)
  7. Demonstration that six months of logs under your control exist now, per high-risk system (Chapter 11)
  8. Litigation hold record following any incident (Chapters 11, 18)

Transparency and oversight

  1. Instructions for use per system: checked against the technical documentation for consistency, against Article 13(3) for completeness, against operating procedures for implementation (Chapters 8, 12)
  2. Oversight design record with rejected alternatives (Chapter 12)
  3. Overseer assignment records with training and authority evidence, reconciled to the current staff list (Chapters 8, 12)
  4. Interface walkthrough record against the Article 14(4) capabilities, override and stop performed (Chapter 12)
  5. Behavioural oversight metrics: override rates, timing distributions, escalations and outcomes (Chapter 12)
  6. Stop-test log, with safe state verified (Chapter 12)
  7. Two-person verification records with competence evidence, for biometric deployments (Chapter 12)

Performance and security

  1. Declared accuracy metrics per system, as stated in the instructions for use (Chapter 13)
  2. Test reports producing the declaration, keyed to the current model version (Chapter 13)
  3. Production monitoring showing the declaration still holds (Chapters 13, 18)
  4. Per-group performance figures wherever the risk file names populations (Chapter 13)
  5. Robustness test plan and results, including failure behaviour and the named manual fallback owner (Chapter 13)
  6. Feedback loop analysis for learning systems (Chapter 13)
  7. AI-specific threat assessment mapped to the Article 15(5) attack families, with incentive reasoning (Chapter 13)
  8. Red-team reports with findings tracked to closure (Chapter 13)
  9. Annex IV Part Seven justification where no harmonised standard was applied: references, date, review trigger (Chapters 13, 16)

Article 50 transparency

  1. Four-duty inventory mapping across all systems (Chapter 14)
  2. Interaction disclosure as the user sees it, per supported language, with the first-interaction design record (Chapter 14)
  3. Marking implementation record: state-of-the-art justification, detection route, version history against the applicable date (Chapter 14)
  4. Synthetic media labelling rule, a sampled published item with its label, and the editorial-responsibility record where the text exemption was relied on (Chapter 14)
  5. Contractual allocation of each Article 50 duty for procured systems (Chapter 14)

The model layer

  1. Model dependency register: every GPAI model relied on, directly or through vendors (Chapter 15)
  2. Per-dependency due-diligence record: Annex XII documentation received, training-content summary reviewed, gaps, questions, answers (Chapter 15)
  3. Contract clauses on version pinning, change notice and data use, per dependency (Chapter 15)
  4. Fine-tuning compute analysis, dated before the training run it assesses (Chapters 2, 15)
  5. Your own downstream documentation pack, plus evidence change notices are delivered (Chapter 15)
  6. Model version recorded per transaction in the logs (Chapters 11, 15)

Conformity assessment

  1. Conformity assessment record: the Annex VI procedure as performed, dated, with the quality management system verification (Chapter 16)
  2. EU declaration of conformity, consistent with the technical documentation, the shipping version and the standards actually applied (Chapter 16)
  3. CE marking as the user encounters it (Chapter 16)
  4. EU database registration, current and matching the declaration (Chapter 16)
  5. Substantial-modification analysis and reassessment, or the pre-determined-changes clause that excused it (Chapter 16)
  6. Signing pack behind the declaration, and the record of who signed: name, role, briefing (Chapter 16)

Audit readiness

  1. Evidence map: obligation to artefact to owner to location, reconciled to the inventory (Chapter 17)
  2. Mock audit reports: findings, owners, closure evidence, retrieval metrics improving between cycles (Chapter 17)
  3. Diary entry for the next mock audit cycle (Chapter 17)

Post-market and incidents

  1. Post-market monitoring plan, guidance once published, with review records at the stated cadence (Chapter 18)
  2. Thresholds, and at least one example of a threshold tripping with what followed (Chapter 18)
  3. Incident procedure: named causal-link owner, operational trigger list (Chapter 18)
  4. Per reported incident: the report, the timeline from first signal to filing against the statutory clocks, the corrective action, the supply-chain notifications (Chapter 18)
  5. Assessment record for any incident assessed and not reported (Chapter 18)
  6. Loop evidence per incident: risk register entry, documentation re-version, instructions-for-use update (Chapter 18)

Enforcement posture

  1. Penalty exposure assessment: systems, tiers, geographies (Chapter 19)
  2. Mitigation posture against the Article 99(7) factors: cooperation procedures, self-reporting thresholds (Chapter 19)
  3. Consistency check between the AI Act file and the GDPR file (Chapters 19, 20)

Multi-framework

  1. Document map per system across all applicable regimes, with the authoritative source named per shared fact (Chapter 20)
  2. Integration evidence where regimes permit merging: combined assessment, merged monitoring, single technical file, with the permitting clause cited (Chapter 20)
  3. Update trail for the last material change: which documents moved, in what order, under whose hand (Chapter 20)
  4. Model-governance-to-AI-Act vocabulary mapping, for financial institutions (Chapter 20)

One hundred and eleven artefacts. If the number alarms you, two important notes. Most artefacts are a page or less, and the sprint in Chapter 21 sequences the first quarter of them into ninety days. 

The alternative to this list is not a shorter list. It is this list, compiled by someone else, attached to a letter.

Sources

  1. Regulation (EU) 2024/1689 of the European Parliament and of the Council of 13 June 2024 (Artificial Intelligence Act), OJ L, 12 July 2024, as amended by Regulation (EU) 2026/1744 of the European Parliament and of the Council of 8 July 2026 (Digital Omnibus on AI), OJ L, 2026/1744, 24.7.2026. (https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng)
  2. European Commission, COM(2025) 836 final, Digital Omnibus on AI, 19 November 2025
  3. DLA Piper, “The Digital AI Omnibus”, knowledge.dlapiper.com, July 2026
  4. Covington & Burling, “EU AI Act Update: Timeline Relief, Targeted Simplification, and New Prohibitions”, insideglobaltech.com, 28 May 2026
  5. Hogan Lovells, “EU legislators agree to delay for high-risk AI rules”, hoganlovells.com, 7 May 2026
  6. ComplianceHub, “What Actually Comes Due on August 2, 2026”, compliancehub.wiki, June 2026
  7. Winston Taylor, “AI Act rules on high-risk AI delayed as AI Digital Omnibus agreed”, winstontaylor.com, 2026
  8. White & Case, “EU agrees Digital Omnibus deal to simplify AI rules”, whitecase.com, 14 May 2026
  9. aiactblog.nl, “The Digital Omnibus and the AI Act: what changes, what now”, June 2026
  10. European Commission, Guidelines on the definition of an artificial intelligence system, C(2025) 924, 6 February 2025
  11. European Commission, Guidelines on prohibited artificial intelligence practices, C(2025) 884, 4 February 2025
  12. Future of Privacy Forum, “Red Lines under the EU AI Act”, fpf.org, 2026
  13. European Commission, Guidelines on obligations for providers of general-purpose AI models, 18 July 2025
  14. Lewis Silkin, “Charting the EU AI Act timeline – key dates from adoption to application”, lewissilkin.com, 16 April 2026. 
  15. Sidley Austin, “EU Lawmakers Reach Provisional Agreement to Delay Key EU AI Act Obligations”, datamatters.sidley.com, 22 June 2026 
  16. Morrison Foerster, “EU Digital Omnibus on AI: What Is in It and What Is Not?”, mofo.com, 1 December 2025.
  17. Cuatrecasas, “Proposal to amend the AI Act: Digital Omnibus”, cuatrecasas.com, February 2026
  18. Gleiss Lutz, “Commission’s digital omnibus proposal to simplify the Artificial Intelligence Act”, gleisslutz.com, December 2025
  19. NicFab, “Digital Omnibus on AI: the European Parliament Rewrites the Commission’s Rules”, nicfab.eu, February 2026
  20. European Parliament, Legislative Observatory, procedure file 2025/0359(COD)
  21. European Parliament Research Service, briefing EPRS_BRI(2026)782651
  22. European Commission, Explanatory notice and template for the public summary of training content for general-purpose AI models, July 2025
  23. European Commission, General-Purpose AI Code of Practice, July 2025
  24. Stibbe, “AI Act reloaded? What the latest AI Act changes mean in practice”, stibbe.com, June 2026
  25. ComplianceHub, “The EU AI Act’s August 2, 2026 Deadline Just Moved — What the Digital Omnibus Actually Changed”, compliancehub.wiki, 10 June 2026.
  26. Lexbeam, “AI Act Penalties and Fines”, lexbeam.com, May 2026
  27. Regulation (EU) 2019/1020 on market surveillance and compliance of products, OJ L 169, 25 June 2019. 
  28. Regulation (EU) 2024/2847 of the European Parliament and of the Council of 23 October 2024 on horizontal cybersecurity requirements for products with digital elements (Cyber Resilience Act), OJ L, 2024/2847, 20.11.2024.