The EU AI Act and content creators: what actually changes in August 2026 (and what doesn't)
The duty to label AI-written text is far narrower than the headlines suggest. And the part that actually affects publishers is somewhere else entirely.
The EU AI Act and content creators: what actually changes in August 2026 (and what doesn't)
From August 2026 most of the European AI regulation applies. In the weeks leading up to it a lot of content circulated saying the same thing: anyone publishing text written with AI will have to declare it, or face multi-million euro fines.
That is not what the law says. The obligation exists, but it is far narrower than the summaries suggest — and it carries an exemption that covers almost everyone frightened by the headline. Meanwhile the part of the regulation that genuinely affects people who publish online is barely mentioned, because it does not make a good headline.
Quick Answer
The AI Act (Regulation (EU) 2024/1689) requires you to disclose AI-generated text only when it is published to inform the public on matters of public interest, and the obligation does not apply where the content has undergone human review and a natural or legal person holds editorial responsibility for it. A marketing guide, a product page or a company post fall outside it. Deepfakes are different: manipulated images, audio and video must always be disclosed, with no editorial exemption. The part that concerns every publisher is elsewhere: providers of general-purpose models must respect the text and data mining opt-out set out in the EU copyright directive, and that opt-out has to be expressed in a machine-readable form. In other words, your
robots.txthas become a legal instrument.
Disclosure
This article is published by the team behind Innova GEO AI Indexer, a WordPress plugin for Generative Engine Optimization. We name our own tool once, at the end, clearly flagged as ours. Everything described here can be done by hand, for free, with a text editor.
Legal note
We are not a law firm and this is not legal advice. It is a technical reading written by people who build tools for publishers, meant to help you ask a professional the right questions. If you run a news outlet, handle sensitive data or work in a regulated sector, go and ask them: we are not the ones who pay for getting this wrong.
The dates, in order
The regulation entered into force on 1 August 2024, but its parts become applicable at different moments. That is why an article announcing that the AI Act "comes into force" reappears every six months: these are successive deadlines of the same text.
| Date | What becomes applicable | Does it concern a content creator? |
|---|---|---|
| 2 February 2025 | Prohibited practices and AI literacy obligations | Marginally |
| 2 August 2025 | General-purpose AI (GPAI) obligations, governance, penalties | Indirectly — see the opt-out section |
| 2 August 2026 | General application, including transparency obligations and Annex III high-risk systems | Yes |
| 2 August 2027 | High-risk systems embedded in already-regulated products | No |
The date that matters for publishers is the middle one: August 2026. From then the transparency obligations can be enforced.
1. The obligation to disclose AI-generated text
This is Article 50, and it causes the most anxiety. It is worth reading for what it says rather than for how it gets summarised.
The duty to disclose that a text was generated or manipulated by AI applies when both of these are true:
the text is published in order to inform the public;
the subject is a matter of public interest.
And it does not apply where the content has undergone human review or editorial control and a natural or legal person holds editorial responsibility for its publication.
What that means in practice
A guide to choosing a CRM, a product page, a company newsletter, a LinkedIn post about your work: none of these inform the public on matters of public interest in the sense the regulation means. They are commercial communication or sector explainers. They fall outside.
A piece commenting on a political decision, reconstructing a news event, or analysing health or environmental data for a general audience: that is inside the perimeter. If AI produced it and nobody reviewed it while taking responsibility, it has to be disclosed.
The point almost nobody reports: the human-review exemption hollows out the obligation for anyone working seriously. If an editor reads it, corrects it and signs off, the duty does not bite — not because you hid something, but because at that point a person is answerable for that text, which is exactly what the regulation is trying to secure.
The other side of it
The exemption is not a stamp you apply to yourself. It assumes the review actually happened and that someone answers for the content. If your process is "generate twenty articles, publish them in bulk, nobody reads them", the exemption does not cover you — and your bigger problem is not the AI Act, it is that you are publishing things nobody has read.
2. Deepfakes follow a different rule
For images, audio and video the logic changes. Manipulated content resembling real people, places or events closely enough to be mistaken for authentic must be disclosed as artificial, and here there is no editorial-responsibility exemption.
There is an accommodation for artistic, satirical and fictional works: in those cases the disclosure can be made in a way that does not spoil the piece — a card in the credits rather than a label across the image.
For anyone producing visual content the operational line is clear: a retouched product photo is not a deepfake; an image showing a real person saying things they never said is.
3. What really affects publishers: the machine-readable opt-out
This is the section the summaries almost always skip, and in our view it is the one that changes the most for people who publish.
Anyone placing a general-purpose model on the European market must put in place a policy to comply with Union copyright law. Among the constraints expressly referenced is the text and data mining opt-out introduced by the 2019 copyright directive: rightsholders may reserve the use of their works for text and data mining, and when they do, that reservation must be respected — including by those training models.
The technical point is in the form. For content made available online, the reservation must be expressed in a machine-readable way. A line in your terms of use written in plain English is not enough.
What counts as machine-readable
robots.txt is the most widespread and universally understood mechanism, and the one major model providers state they honour. More structured proposals exist, but today the file at the root of your domain is the instrument that works.
Which leads to a counter-intuitive consequence: your robots.txt has stopped being a technical preference and become the place where you exercise — or fail to exercise — a right. Write nothing in it and you have reserved nothing.
4. The tension nobody resolves for you
Here is the uncomfortable part, because it is the real fork in the road and you will not find it in the enthusiastic summaries.
Blocking AI crawlers to reserve your rights and wanting to be cited by AI assistants are two goals in tension. Not entirely incompatible, but in tension. And the AI Act does not tell you which to choose: it hands you the instrument to exercise a choice that remains yours.
What makes the choice manageable is that these bots do not all do the same job.
| Bot type | What it does | If you block it |
|---|---|---|
| Training crawler (e.g. GPTBot, ClaudeBot, CCBot) | Collects text to train models | You exercise the TDM reservation. You do not lose citations in answers. |
| Retrieval / search crawler (e.g. OAI-SearchBot) | Feeds the index the assistant draws on when answering | You drop out of cited answers. This is where visibility is lost. |
| User-triggered fetch (e.g. ChatGPT-User) | Opens a page because a user asked for it | The assistant cannot open your link when someone requests it. |
The coherent position for most publishers is: reserve training, stay open to retrieval. It is not the only legitimate position — a publication that licenses its archives may decide otherwise, with excellent reasons. But it is the position that maximises visibility without giving up the right.
What is definitely wrong is pasting a block copied from a 2024 article without looking at what is in it. Those blocks lumped together bots from different categories, and the sites that applied them removed themselves from assistant answers while believing they were protecting themselves from training.
5. What the AI Act does not say
Three claims in circulation that are not in the text:
"Every piece of AI-written content must be labelled"
No. The duty is limited to informative text on matters of public interest, with the human-review exemption. The rest of online communication does not fall under this obligation — which does not make disclosing it wrong, only that it is not imposed here.
"If I use AI to write, I become an AI system provider"
No. Someone using a tool in their own work is a deployer, not a provider. The heavy obligations — conformity assessments, technical documentation, marking of outputs — sit with whoever builds the model and places it on the market. Write an article with an assistant and those duties are not yours.
"I risk multi-million fines on my blog"
The regulation's highest penalties concern prohibited practices, which have nothing to do with producing content. Breaches of transparency obligations sit in a lower band, and for SMEs the regulation expressly applies the more favourable of the percentage-of-turnover and fixed-amount caps. The concrete risk for an average editorial site is not the fine: it is the loss of credibility from publishing unverified content.
Practical checklist
Five things to do, ordered by how much they matter and how much they cost.
Open your
robots.txtand read it. If it contains a block copied years ago, check which user agents you are blocking and which category they belong to. This is the most important item on the list and takes ten minutes.Decide your position on training and write it into the file explicitly. "Allow everything" is a position too — but it should be a choice, not an oversight.
Look at what you actually publish. If your output includes informative pieces on matters of public interest, define who reviews them and who answers for them. This is not paperwork: it is the exemption that keeps you outside the obligation.
If you produce realistic images or video of real people, prepare the disclosure. The editorial exemption does not exist here.
Write the process down. Who generates, who reviews, who publishes. You do not need a quality system: you need to be able to answer "who checked this piece" with a name.
Where our tool comes in
We have not named a product until now because there was no need: everything above can be done with a text editor and half an hour of attention.
If you run a WordPress site, though, and item 1 on the checklist made you wince — because you know you have a robots.txt nobody has looked at in years, or because you do not want to memorise which user agent does what — Innova GEO AI Indexer is the plugin we build, and it is free on the official WordPress repository.
For the subject of this article it does two things:
it shows the known AI crawlers inside WordPress, one by one, with what each is for and what you lose by blocking it written next to it, and generates the matching
robots.txtrules. All are allowed until you decide otherwise: the choice is yours, made with the facts in front of you;it keeps a log of the AI generations made through the plugin — useful when you need to reconstruct who generated what, and when.
What it does not do, and no plugin can do for you: decide your position on training, establish who takes editorial responsibility for your content, or automatically label your text as AI-generated. Those are decisions, not features.
Frequently asked questions
Do I have to put "written with AI" on all my articles?
Not as an AI Act obligation, unless the text is informative content on a matter of public interest without human review. You may do it as a transparency choice towards readers, which is a different and perfectly good reason — but it is your editorial decision, not a compliance step.
If I block every AI bot, am I safe?
You are safe with respect to training, and invisible to assistants. It makes sense if your business is selling access to content, much less if you live on traffic and visibility. The regulation does not require you to block: it gives you the means to do so if you want.
My site is outside the EU — does this concern me?
The regulation has extraterritorial reach tied to the European market, and primarily concerns providers and deployers of AI systems whose output is used in the Union. For a non-EU publisher the practical question is different: the model providers whose respect for your opt-out you actually want do operate in Europe, so they look at your robots.txt anyway.
Does the opt-out cover what I have already published?
The reservation takes effect when you express it: it does not retroactively erase what has already been collected. Which is why postponing the decision carries a cost you cannot recover.
Is an image generated from scratch a deepfake?
Not if it does not depict real people, places or events in a way that could be mistaken for authentic. An abstract illustration or an obviously imaginary subject does not fall under the deepfake disclosure duty — the technical marking of outputs still applies, but that is the responsibility of whoever supplies the generator, not yours.
Where do I read the original text?
Regulation (EU) 2024/1689 is published in the Official Journal of the European Union and is freely available on EUR-Lex in every official language. If even one part of this article touches you closely, it is worth the half hour to read the cited articles in the official version rather than trusting a summary — including ours.
Frequently asked questions
Does the EU AI Act require me to label all AI-written content on my website? +
No. Article 50 of the AI Act limits the disclosure duty to informative text on matters of public interest produced without human editorial review. Commercial content, marketing copy, product pages, and company posts fall outside this obligation. You may disclose AI use voluntarily, but the Act does not mandate it for most online content.
What is the human-review exemption in the EU AI Act for content creators? +
If a human editor reads, corrects, and takes editorial responsibility for AI-generated text before publication, the obligation to disclose AI authorship does not apply. The exemption reflects the regulation's goal: ensuring a person is accountable for published content, not banning AI assistance in content creation.
Does the EU AI Act apply to deepfakes differently from text? +
Yes. For manipulated images, audio, and video depicting real people, places, or events in a way that could be mistaken for authentic, disclosure is always required. Unlike text, there is no editorial-responsibility exemption for deepfakes, though artistic and satirical works may carry a less intrusive disclosure.
How does robots.txt relate to the EU AI Act? +
The AI Act requires GPAI model providers to respect the text-and-data-mining opt-out established by the 2019 EU Copyright Directive. That reservation must be expressed in machine-readable form. The robots.txt file at your domain root is the primary recognised mechanism, making it a legal instrument rather than just a technical preference.
When does the EU AI Act transparency obligation for publishers come into force? +
The transparency obligations under Article 50, including disclosure of AI-generated informative content, become enforceable on 2 August 2026. General-purpose AI obligations and penalties applied earlier, from 2 August 2025. The regulation entered into force on 1 August 2024 but applies in staggered stages.
If I block all AI crawlers in robots.txt, am I fully protected under the EU AI Act? +
Blocking all AI crawlers protects your text-and-data-mining rights regarding training, but it also removes your content from AI assistant retrieval indexes, reducing your visibility in AI-generated answers. The regulation does not require blocking; it gives you the instrument to choose. Training crawlers and retrieval crawlers serve different purposes and can be managed separately.
Sources
- Regulation (EU) 2024/1689 – EU AI Act, Official Journal of the European Union (EUR-Lex)
- Directive (EU) 2019/790 on Copyright in the Digital Single Market – Text and Data Mining provisions (EUR-Lex)
- EU AI Act – European Commission official overview
- AI Act transparency obligations for deployers – European Parliament briefing