Generative AI has moved from experiment to infrastructure. Marketing teams draft campaigns with it, developers write code with it, and consultants produce first drafts of reports with it. The legal questions that follow are rarely asked until something goes wrong: who owns the output, can it be protected, and what happens if the model reproduced someone else’s work along the way?
Why copyright is the sharpest edge of AI adoption
Copyright arises automatically. There is no register to check, no application to file, and no fee to pay – which means there is also no easy way to confirm that a piece of AI-generated text, image or code is clear of third-party rights. For a business, that creates a quiet accumulation of risk across thousands of everyday outputs.
Two separate questions need to be kept apart. The first concerns the input side: was it lawful to train the model on protected works? The second concerns the output side: can what the model produces be protected, and might it infringe? Most commercial disputes turn on the second, but the first shapes the contracts you sign with your AI vendors.
The input side: text and data mining in EU law
EU copyright law does not prohibit training outright. The Directive on Copyright in the Digital Single Market introduced two exceptions for text and data mining. One covers research organisations and cultural heritage institutions carrying out mining for scientific research. The other is broader and available to anyone, including commercial developers – but it applies only where the rightsholder has not expressly reserved the use.
That reservation is the mechanism at the centre of the debate. For content made publicly available online, the reservation must be expressed in machine-readable form. In practice this has meant robots.txt directives, metadata signals and emerging rights-reservation protocols. If a rightsholder has opted out in a way a crawler can read, the commercial exception is not available.
The AI Act reinforces this. Providers of general-purpose AI models must put in place a policy to comply with EU copyright law, including identifying and respecting rights reservations made under the text and data mining exception. They must also publish a sufficiently detailed summary of the content used for training, using a template published by the European Commission. That obligation took effect on 2 August 2025 for new models, with models placed on the market before that date having until 2 August 2027. From 2 August 2026 the Commission gains enforcement powers, with fines of up to €15 million or 3% of worldwide turnover for GPAI providers.
The output side: can AI-generated work be protected?
Under EU and Swedish copyright law, protection requires a work that is the author’s own intellectual creation reflecting free and creative choices. The Court of Justice has consistently anchored protection in human creativity. Material generated autonomously by a model, on a bare prompt, is very difficult to fit within that requirement.
This does not mean AI-assisted work is unprotectable. Where a human contributes meaningful creative choices – the selection, arrangement, editing and refinement that shape the final result – the human contribution can attract protection. The practical consequence is a spectrum: the more the output is a product of substantive human authorship, the stronger the position; the closer it is to raw model output, the weaker.
For businesses, the commercial implication is direct. Content you cannot claim copyright in is content a competitor may be free to copy. If a logo, a campaign visual or a body of documentation matters to your brand, the process by which it was created matters too.
Infringement risk in outputs
Models can reproduce material closely resembling their training data, particularly for distinctive styles, well-known characters, code snippets and stock imagery. Infringement does not require intent, and it does not require you to know where the model got it. If an output reproduces a substantial part of a protected work, the business publishing it carries exposure.
Code deserves particular attention. Open source licences carry conditions – attribution, and in the case of copyleft licences, obligations that can extend to derivative works. An AI coding assistant that emits a licensed snippet into a proprietary codebase can create a licence compliance problem that surfaces years later during due diligence.
Practical example: the campaign that could not be registered
A Swedish consumer brand generates a set of illustrations with an image model, selects one with minimal editing, and rolls it out across packaging and advertising. Six months later a competitor launches with a strikingly similar visual identity.
The brand discovers two problems at once. It has difficulty asserting copyright in the illustration, because the human creative contribution was limited to choosing between generated options. And its agency contract assigned “all intellectual property in the deliverables” – a clause that transfers whatever rights exist, but cannot create rights that never arose. Had a designer used the generated image as a starting point and made substantive creative choices, both the protection and the negotiating position would have been materially stronger.
Common mistakes companies make
Assuming vendor indemnities solve the problem. Several providers offer copyright indemnities, but they are typically conditional on using the service as instructed, keeping safety filters enabled, and not supplying infringing input. Read the conditions before relying on the promise.
Feeding confidential material into public tools. Uploading customer data, draft agreements or unpublished technical information into a consumer-grade assistant can breach confidentiality obligations and, if the information is a trade secret, destroy the reasonable steps required to keep it protected.
Treating “we own the output” in the terms of service as ownership. A provider assigning its rights to you says nothing about third-party rights in the same material.
Having no record of how content was made. When a dispute arises, the ability to show human creative input is evidential. Few organisations keep any trace of it.
Overlooking the transparency duty to customers. Where AI-generated content is presented as human-made in a way that could mislead consumers, marketing law and the AI Act’s transparency obligations both become relevant.
Recommended actions
Adopt a written AI use policy that states which tools are approved, what may and may not be entered into them, and where outputs require human review before publication. A policy nobody has read is worth little, so pair it with short, role-specific training.
Tier your content. Low-stakes internal drafts need light governance. Anything that becomes brand identity, product, published code or a contractual deliverable needs documented human authorship and a review step.
Keep provenance records for material that matters: the prompts, the iterations, and what a person changed. Run similarity checks on visual assets and licence scanning on AI-assisted code before release.
Review your supplier and customer contracts. On the supplier side, look for training-data warranties, indemnity scope and carve-outs, and confirmation of AI Act compliance. On the customer side, be careful about warranting the originality of deliverables you produced with AI assistance.
Finally, decide your own position as a rightsholder. If your website content has commercial value, consider whether to reserve your rights against text and data mining in machine-readable form.
Frequently asked questions
Who owns content created by our employees using AI?
Where the output attracts copyright through human creative contribution, ordinary rules on employee-created works and your employment contracts determine ownership. Where no copyright arises, there is nothing to own or transfer – which is why contractual assignment clauses cannot fix the underlying issue.
Can we stop AI companies from training on our website?
You can reserve your rights under the text and data mining exception, and for publicly available online content that reservation should be machine-readable. Providers of general-purpose models are required under the AI Act to have a policy for identifying and respecting such reservations.
Do we have to tell customers that content is AI-generated?
The AI Act imposes transparency duties for certain AI-generated or manipulated content, including deep fakes, and requires machine-readable marking of synthetic output. Separately, consumer protection law prohibits misleading commercial practices. The safe position for customer-facing material is disclosure where the origin would matter to the audience.
Is using AI to summarise copyrighted documents an infringement?
Summarising for internal use generally sits differently from reproducing and publishing protected expression, but the analysis depends on what is copied and what is done with it. Uploading third-party material can also breach the confidentiality or licence terms under which you hold it.
Conclusion
Generative AI does not sit outside copyright law – it sits awkwardly inside it. The workable approach is neither prohibition nor indifference, but proportionate governance: know which tools are in use, keep confidential material out of them, insist on genuine human authorship where the output matters commercially, and hold your vendors to the transparency obligations the AI Act now places on them.
Lawgent advises companies on AI use policies, supplier and customer contract terms, intellectual property strategy for AI-assisted output, and compliance with the AI Act’s transparency and copyright provisions. Contact us to review how your organisation is using generative AI today.