TL;DR
I am a software engineer in London, in the early stages of forming an AI startup.
Different copyright exceptions can leave a UK commercial developer unable to rely on datasets widely used in other jurisdictions.
Project Gutenberg, for example, does not provide a reliable UK public-domain flag, so its eBooks need separate UK verification.
Publicly and permissively licensed text can still contain personal data, and even screening it to exclude sensitive material is processing.
A curated professional and official-text corpus may still be defensible through legitimate interests, minimisation, pseudonymisation, transparency and a DPIA. The hardest unresolved point is genuinely incidental Article 9 and Article 10 intake.
1. What I am building, and the two kinds of provenance
I am a software engineer in London, in the very early stages of forming an AI startup. Our first product is a classifier model that labels text as human-written or AI-generated, similar to tools such as GPTZero and Pangram.
Open research and increasingly mature tooling make model experimentation accessible. For this project, however, architecture is not the first or hardest constraint. The difficult asset is a corpus whose labels and reuse rights can both be evidenced.
For every document, or every demonstrably homogeneous source component, I need to establish two independent facts: rights provenance (am I allowed to use this text commercially?) and generation provenance (do I know whether and how a human or a model produced it?). A dataset card establishes neither by itself.
Generation evidence comes in degrees. The strongest evidence I can routinely obtain is a first-party account from the publisher or creator explaining how the text was produced. At the other end, I exclude material with unknown or inseparably mixed provenance. Publication before the generative-AI era is strong supporting evidence, though not conclusive proof, since template-generated text predates large language models by decades. The AI-generated half of the corpus needs the same care: I record the model, version, generation date, prompt lineage and any human editing.
The first admission screen is a matrix: strong rights evidence and strong generation evidence are both necessary. Failure on either means rejection or quarantine; passing both only advances the material to the access, privacy, safety, attribution and reproducibility gates. A mixed release can pass where human and model material is separated by a documented deterministic rule. Unknown or inseparably mixed provenance is the exclusion class.
I began with the naive assumption encouraged by the research literature. Papers often describe acquisition at a high level: “we collect a large dataset of 300K human-written stories from an online forum”. Fan et al. (2018) report 303,358. Surely I could download the published datasets and get on with the machine learning? Under UK law, that single question breaks into several legal and evidential ones.
The system should present materially less risk to people in the training corpus than a generative model. It returns labels only; identifiers are meant to be removed before training, and it has no ordinary way to reproduce anyone’s prose. That remains a design claim until the intake controls and model-privacy tests described below are in place. The greater practical risk may fall on the people whose writing the deployed classifier later scores. I return to them in section 7 because that is where I think the product’s moral weight lies.
2. A dataset licence may not license its contents
Copyright is the first complication. A “licensed dataset” can bundle together four separate layers:
contractual permission to access the platform or API the text came from;
copyright in the individual texts;
copyright in the selection or arrangement and any sui generis database right;
the licence covering the repository code, the metadata, and the packaged database as an artefact.
A GitHub repository’s MIT licence, or a Hugging Face “apache-2.0” tag, may govern only one of those layers, sometimes only the repository code, metadata or packaged artefact. The Government’s report on copyright and artificial intelligence (March 2026) distinguishes between licensed datasets, whose underlying rights have been obtained and assembled, and collections whose component works remain separately protected. It also notes that downloading and preprocessing in the UK can involve fresh acts of copying even when the dataset was assembled abroad. My ingestion may therefore involve restricted acts that need their own licence or statutory justification. The upstream collector’s position does not automatically shelter mine.
The problem is widespread. The Data Provenance Initiative audited 1,858 text fine-tuning datasets across 44 collections. Around 70% of dataset entries on major hosting platforms did not specify a licence. Of the Hugging Face licence annotations it examined, 66% fell into a different use category from the researchers’ determination of the underlying licence. The label on the tin is often not the licence of the contents.
WritingPrompts is a useful example. The Fan et al. corpus contains 303,358 stories collected from a Reddit forum through Reddit’s official API and distributed through the research project associated with an academic paper. Under Reddit’s current User Agreement, users retain ownership of their posts while giving Reddit a broad licence that includes AI uses. Reddit’s current Developer Terms grant no right to train AI or machine-learning models on user content without the applicable rightsholders’ express permission. Those current terms do not retrospectively determine the researchers’ position in 2018 or show that the original collection was unlawful. But an academic release does not prove that a commercial company in London may now train a product on the 303,358 stories Reddit users contributed. Nor does it establish that the dataset creators had authority to license every story for that use.
Section 29A of the CDPA 1988, the UK’s text-and-data-mining exception, permits copies for computational analysis only with lawful access, for the sole purpose of non-commercial research, and with sufficient acknowledgement where practicable. Corporate status alone is not disqualifying; purpose is what matters. I cannot safely rely on the exception for commercial product training.
Other jurisdictions take different approaches, though they do not authorise restricted acts carried out in the UK. Japan’s Article 30-4 extends beyond non-commercial actors under a “non-enjoyment” framework, subject to a proviso against unreasonable prejudice to rightsholders; the Agency for Cultural Affairs gives recording works in an AI-training database as an example. Singapore’s ss.243-244 permit computational data analysis subject to lawful access and restrictions on use. The EU’s DSM Directive Article 4 provides an opt-out TDM exception, with online reservations expected to be machine-readable. The UK declined its broad 2022 proposal, and the March 2026 report says a broad opt-out exception is no longer preferred, although targeted reforms remain under consideration. UK law therefore provides no general commercial TDM exception for those acts today.
There is no UK fair-use route either. In Bartz v. Anthropic, the June 2025 summary-judgment order held that using the books at issue for the particular training use was “exceedingly transformative” fair use. The court analysed the acquisition and retention of shadow-library copies as a separate use and denied Anthropic summary judgment on it. The $1.5bn settlement later resolved the litigation risk without a merits judgment pricing acquisition alone.1 The useful lesson is narrow: acquisition must be audited separately from training. A UK analysis cannot simply import a fact-specific US district-court ruling.
Licensed human text is not vanishingly scarce. Common Pile v0.1 assembled eight terabytes from 30 public-domain and openly licensed source families and trained seven-billion-parameter models competitive with comparable budget-matched models trained on conventionally sourced data. The Open Parliament Licence makes the copyright and database rights in a large body of Parliamentary material available for reuse, but it expressly excludes personal data and third-party rights. Solving copyright does not settle data protection. The European Commission’s DGT translation memory carries an official reuse grant. My project’s admission policy accepts express document-level licences, UK-verified public domain, official open licences and direct commissioning. That is my conservative risk threshold, not an exhaustive statement of what the law permits. Even the public-domain route needs per-work, per-country verification. Project Gutenberg, for example, is “entirely based in the US”, applies US copyright law and warns that “not all items that are public domain in the US are public domain in other countries”. It does not supply a reliable UK-public-domain flag, so candidate works need separate UK verification.
3. Public text can still be personal data
Data-protection law comes into play whenever text contains information relating to an identified or identifiable living individual. It does not cover the deceased, information solely about a legal person, or people who genuinely cannot be identified. Information about a sole trader, director, employee or named company contact can still be personal data. Within that boundary, Article 4(2) defines “processing” so broadly that downloading a dataset to inspect it is already processing whatever personal data it contains. Because the sources are heterogeneous and unstructured, I cannot reliably determine every category of personal data in each document without first inspecting it.
The ICO rules out the most obvious escape. Its tackling-misconceptions page states directly: “The ‘incidental’ or ‘agnostic’ processing of personal data still constitutes processing of personal data”, regardless of what the developer intended or cared about. Its web-scraping analysis adds that intention is irrelevant to whether existing content falls within Article 9’s special categories. The ICO treats inferential data differently. Where information does not itself clearly reveal a special category, Article 9 is engaged if the controller intends to infer that characteristic or treat the person differently on that basis.
Public availability is not an exemption. There is no automatic transparency exception for information in the public domain (ICO right-to-be-informed guidance), and even a very open copyright licence does not waive anyone’s data-subject rights. The Open Parliament Licence’s express carve-out of personal data states the position plainly.
There are two important limits. First, not every mention of a person engages Article 9. The material must reveal a sufficiently definite protected characteristic of an identifiable living person. A news article naming an MP does not, by itself, reveal a personal political opinion. A report that identifies a living person’s medical diagnosis is different. Second, public availability and professional context can affect reasonable expectations and impact within a legitimate-interests balance, even though they are not exemptions. That is why professionally bylined and official text is my strongest starting cohort. Section 5 explains the proposed route.
Obvious statutory candidates do not automatically solve the sensitive-data problem. Article 9(2)(e) generally requires the data subject to have manifestly made the information public, not merely a journalist or public body to have published it. Article 9(2)(j) and DPA 2018 Schedule 1 paragraph 4 may cover some commercial technological research where it can reasonably be described as scientific, but necessity, safeguards and a qualifying public interest remain required; Schedule 1 paragraph 32 provides a similarly narrow subject-publication route for offence data. Whether any of those conditions covers fleeting intake performed solely to identify and discard unwanted material remains part of the uncertainty.
4. The filtering paradox and the emerging European answer
The strongest sceptical argument against this position begins with ordinary web browsing.
When an employee’s browser renders a news page about a living person’s diagnosis, the device downloads, copies, parses and displays health data. On the ICO’s logic, that is processing. Yet organisations do not ordinarily conduct a page-specific DPIA merely because an employee might encounter personal data while browsing public pages. Does that make the claim that “downloading is processing” too broad to be useful?
The analogy’s legitimate insight is that a workable account of data-protection law must distinguish ubiquitous transient technical handling from systematic corpus construction. But the distinction is available without pretending the browser never processed the text: purpose, scale, systematicity, persistence, aggregation, reuse, identifiability and risk all differ between a page render and a training corpus. Article 35 is risk-triggered, not universal. Fashion ID (C-40/17) supports the narrower proposition doing the real work: responsibility attaches to the operations for which an actor determines the purposes and means. Whether, and for which stages, an employee or employer is controller for ordinary browsing is a further question. I do not determine the publisher’s original disclosure, a search engine’s indexing or every independent browser operation; I do determine my corpus acquisition, screening, retention and reuse.
The claim still holds, but it creates a real paradox for anyone trying to comply carefully. A privacy screen must read a document to discover Article 9 special-category material or Article 10 criminal-offence material. Those are different legal gateways and require separate analysis. Inspection for the purpose of deletion is already processing. Retaining candidate text indefinitely without screening is also processing and is usually harder to justify, since it abandons minimisation for no benefit. Both the most protective and the least protective available acts count as processing, and the protective act is the one that touches the sensitive data. Copyright law at least has an express, narrowly conditioned temporary-copy rule for one version of this problem: s.28A CDPA. UK GDPR has no equivalent statutory safe harbour for transient intake.
European regulators have recently begun addressing this problem.
The CNIL, in its adopted AI guidance (lawful-basis and web-scraping sheets), addresses the case of a controller that has implemented exclusion measures but nevertheless “processes incidentally and residually sensitive data that it had not sought to collect”: such processing “is not considered illegal”, drawing on the CJEU’s reasoning in GC and Others (C-136/17). The CNIL couples that conclusion with preventive collection criteria, source exclusions, prompt deletion and safeguards including pseudonymisation. This is adopted French guidance, not a draft.
The EDPB, in draft Guidelines 03/2026 on web scraping in the generative-AI context (adopted 7 July 2026 for public consultation, open until 30 October 2026), acknowledges at §§67-77 that a controller often cannot know whether scraping collects Article 9 data “until after the data has been scraped”. It proposes applying the GC and Others “responsibilities, powers and capabilities” analysis case by case to special-category collection that is incidental and residual, never intentional. The later paragraphs make that analysis depend on factual similarity to the search-engine situation, an inability reasonably to prevent collection in advance, immediate deletion, source exclusions, accountability and continuing testing. The draft expressly says that this “should not be seen as a general exemption from the requirements in Articles 9 and 10”.
The ICO has said that incidental processing still counts and warns against casually transplanting search-engine case law into generative AI. GC and Others is search-engine case law. The ICO has not explained how the judgment applies to a materially different UK intake-and-exclusion pipeline.
The position has four parts. Under UK statute and current ICO guidance, intake is processing; incidental processing remains within UK GDPR; Articles 9 and 10 require their own gateways; and deletion is a safeguard rather than a lawful basis. As assimilated EU case law, GC and Others remains relevant to the corresponding UK GDPR provisions, subject to its factual and legal relevance, subsequent UK amendments and the appellate departure regime. Under s.6 of the European Union (Withdrawal) Act 2018, relevant pre-completion CJEU decisions form part of assimilated EU case law. In a non-binding extrapolation, the CNIL’s adopted position and the EDPB’s draft extend the judgment’s search-engine reasoning to carefully bounded incidental intake. The remaining UK question is whether that reasoning reaches a licensed, curated, non-generative intake-and-exclusion pipeline like mine, and whether any comparable route exists for Article 10. That gives the ICO a more concrete question than “nobody has addressed this”. I ask it in section 9.
In the meantime, my project’s proposed design mirrors many of the controls the CNIL and EDPB describe: pre-collection source criteria and exclusions, a quarantined intake zone, automated screening, immediate deletion of what should not be there, and pseudonymisation of what remains. That does not establish that the legal threshold or the required factual similarity to search-engine processing is met; I need to know the UK answer before the intake stage ever runs against personal-data sources.
5. What defensible processing would actually look like
I think processing the strongest cohort is probably defensible if it is carefully governed. The proposed system would work as follows.
The corpus lifecycle is a sequence of stages, each with a different purpose, exposure and risk: source catalogue → rights review → candidate intake → quarantine and privacy screen → exclusion or pseudonymisation → admitted corpus → training and evaluation → model privacy test → release → deployment. “The dataset” is not one legal object. Candidate intake is an unclassified zone; no raw candidate material is admitted to the training corpus or exposed to training jobs.
The most plausible route is legitimate interests plus minimisation. A commercial interest can be legitimate (ICO legitimate-interests guidance); the three tests are purpose, necessity and balancing. My working hypothesis, which still needs testing, is that authentic human prose containing the minimum incidental personal data is necessary to train and test the classifier. Names and sensitive facts are not necessary model features. I would process them transiently only for screening, provenance and rights handling. Before relying on necessity, I intend to record comparative experiments using less intrusive alternatives: a corpus with no personal data, synthetic text, public-domain-only text and more heavily sanitised text. The aim is to measure whether any of them works equally well. For the balancing test, the strongest starting cohort is deliberately published professional or official writing; incidental mentions of other people need a separate assessment. The model is designed to infer and decide nothing about people in the corpus. It emits a label and model score, never prose. Objection, erasure and source-exclusion routes must work before the gate opens. I will not know whether the trained weights or outputs reveal anything about people in the corpus until the model-specific tests in section 6 have run.
At corpus scale, transparency is governed by Article 14(5)(e), (6) and (7) as amended by DUAA 2025 s.77. Non-notification requires a documented proportionality assessment and protective measures that include public information. The assessment must distinguish materially different cohorts instead of treating corpus scale as a blanket answer. Not notifying a readily contactable author with a current professional byline requires stronger justification than not notifying an incidental person identified only indirectly in a decade-old report, for whom no practicable contact route exists. Where the assessment supports the exception, I propose a prominent training-data notice and a public registry naming sources, versions and acquisition dates where feasible. There would also be a plain account of the admission and sanitisation stages, a form for objections, erasure and source exclusions, a search route by URL, byline or text, and a published retention and release policy. Individual notice remains the default unless the assessment justifies the exception for that cohort.
Pseudonymisation is not anonymisation. The provenance vault allows me to evidence rights and honour rights requests, which means I can reconnect sanitised text to its source. The internal store therefore remains pseudonymised personal data (ICO pseudonymisation guidance). Replacing names does not end the legal analysis. It is a safeguard used for minimisation, security, balancing and residual-risk acceptance.
This project requires a DPIA. The proposed system combines innovative AI with indirect, potentially large-scale and partly invisible processing. The assessment must determine whether the implemented controls reduce residual risk below the Article 36 prior-consultation threshold.
Everything above is still design rather than operation. The DPIA and retention schedule are drafts that describe conditions precedent. The privacy gate stays closed until I have documented an Article 9/10 route and gathered implementation and test evidence for the controls. No corpus-scale ingestion of body text, sanitisation or model training on personal-data sources has begun. Limited source due diligence and metadata processing is governed separately.
A current hosting page cannot be assumed to describe an older derivative accurately. The operational rule is to hash the acquired version, retain the dataset card and licence text as they stood at acquisition, record component versions, and monitor upstream withdrawals and rights notices. Aggregate datasets need rights review at both component and version level because a dated derivative can retain a component that its upstream later withdrew.inheritance
The principle behind these controls is that compliance is not automatically inherited. Provenance, licences and other permissions may pass downstream where their terms and legal effect allow, but an upstream actor’s lawful processing does not establish the next controller’s compliance. A US academic’s lawful collection does not, by itself, supply a UK commercial user’s permissions. A gap downstream does not imply that the academic acted unlawfully.
6. Rights requests and model weights
What happens when someone asks what I hold about them?
Article 15 gives a right to confirmation, a copy and supplementary information. After the DUAA, UK law expressly requires reasonable and proportionate searches (ICO summary). A requester cannot prescribe unlimited searches detached from personal data that may reasonably be located. The question is whether particular data is personal data within scope and can be found through a reasonable and proportionate search.
Article 11, which relieves a controller that cannot identify data subjects, may offer this project less relief than it first appears. I would rather acknowledge that now than discover it in correspondence. A provenance vault, source identifiers and consistent replacement tokens may mean I am in a position to identify. A controller cannot ignore the means available to it and then invoke Article 11. The ICO’s engineering-individual-rights analysis says organisations should not view it as a way to avoid obligations.
The cost of honouring rights depends as much on corpus organisation as on scale. Pseudonymisation neither removes information rights nor makes searches impossible. A well-indexed vault with a consistent replacement scheme can make authorised searches and source suppression easier, while limiting identification during ordinary training operations. The same infrastructure that weakens my Article 11 position makes it easier to honour Articles 15 to 21. I think that is the right trade.
The weights raise a separate question. The Hamburg data-protection authority’s 2024 discussion paper, expressly a discussion paper, takes the position that storing an LLM does not amount to storing personal data in the model merely because personal data contributed to training. EDPB Opinion 28/2024 takes a case-by-case approach. It treats a model as anonymous only when the likelihood of extracting training-subject data or eliciting it through queries is “insignificant”, considering all means reasonably likely to be used. The ICO’s own guidance also rejects an automatic assumption that models are not personal data. It points to model inversion, membership inference and architectures that retain examples by design. A label-only classifier may have a stronger claim to anonymity than a generative model because it has no ordinary output channel for prose. But model scores can leak membership information, and an API has a different attack surface from downloadable weights. The claim means little until tested. My position is therefore a hypothesis with a test plan: membership-inference and extraction testing against the actual model in its actual release context, before I make any claim of anonymity.
7. The risk after training: false accusations
The greater moral risk may come after training. So far I have discussed people whose published prose contributes statistical signal to a label-only model designed to reduce risk to its training subjects. Students, employees, journalists and job applicants face a different, and potentially much larger, risk when the deployed classifier scores their work and falsely identifies it as AI-generated.
Liang et al. found that several widely used detectors systematically misclassified non-native English writing as AI-generated. In their TOEFL-essay sample, the average false-positive rate exceeded 60%, while the comparison sample of native-English writing produced a much lower rate. The study evaluated particular detectors at a particular time, so it does not establish an unavoidable flaw in every detector. It does show why each product needs subgroup testing and cautious deployment.
That leads to several deployment commitments. The output is an uncertain statistical estimate, never proof. Any value presented as a probability must first be shown to be calibrated for the relevant domain and text length. The terms must prohibit its use as the sole basis for disciplinary, employment or educational action. The model should abstain below a minimum text length and report uncertainty as a first-class output. I will test calibration by domain, language background and editing level, publish comprehensible false-positive rates, and warn explicitly about domain shift and human-edited text. Where customers use the output consequentially, affected authors need meaningful review as a product and fairness safeguard. A school, employer or platform using a score about an identifiable author must assess its own lawful basis, fairness, accuracy and transparency obligations. If the processing is likely to result in high risk, it must complete a DPIA. If it makes a legal or similarly significant decision solely through automated processing, Articles 22A–22C require clear information, a route to contest or make representations, and human intervention. A customer’s misuse is not automatically my processing, but foreseeable misuse must shape product design, terms and release decisions.
Bias testing presents a smaller version of the same governance problem. Demonstrating fair performance across groups may require a benchmark containing protected-characteristic or closely correlated data. That benchmark needs its own lawful basis, Article 9 analysis where applicable, minimisation and governance; the answer cannot be to infer sensitive attributes casually from names or writing.
This asymmetry also grounds a proportionality argument I want regulators to hear: the law should distinguish a potentially lower-impact training-time contribution from high-impact deployment. The privacy risk to an official whose byline is removed before training may be much lower than the fairness risk to a student later accused on the strength of a false positive. Weighting scrutiny accordingly is proportionate regulation, not opposition to regulation.
8. Why uncertainty has a fixed cost
My complaint is narrower than these discussions often become. Most of the individual obligations above are reasonable, and a careful engineer would adopt several of them anyway. The problem is that broad duties combined with sparse authority create an uncertainty cost with a large fixed component. Small firms have less capacity to absorb that fixed cost. I am not claiming that they always have the highest total compliance expenditure. I am saying that a legal opinion, provenance audit or delayed experiment consumes a much larger share of a two-person firm’s capital and attention than of a platform’s.
The Government’s March 2026 copyright report supports part of this concern: SMEs and individual developers may find it difficult to take advantage of other countries’ regimes, given the resources needed to research and ensure legal compliance. The ICO’s Data Controller Study 2024 reports that organisations experience a lack of clarity and uncertainty about innovative products. Beyond that, the comparison is my own experience and inference, and I label it as such.
The clearest evidence I can offer is the project’s own numbers, the sort of worked example DSIT’s call requests. As of 19 August 2026, the source catalogue records 688 source or source-family rows examined across nine text classes. Of those, 103 are rejected, mostly because the underlying rights are unclear or the generation evidence is inadequate. Another 391 sit in rights-review, 53 in access-review, and 141 have passed evidence capture as candidates. Zero have cleared every gate to become “approved” because the privacy gate has not opened. A parallel audit reviewed 106 PDFs about human-text datasets. It rejected 41 releases outright; 63 remained in rights review, including one split-component disposition; and two were not text-dataset papers. Those figures represent months of my time spent on rights and provenance rather than on models. They also represent a large volume of technically useful text unavailable to the project at present because I cannot establish, within the project’s admission policy, a sufficiently evidenced lawful route at a price a two-person company can pay.
The scope point is often overstated. Processing carried out in the context of my London establishment is straightforwardly within UK GDPR. A non-UK controller without a UK establishment falls within Article 3(2) only when its processing relates to offering goods or services to people in the UK or monitoring their behaviour. The downstream issue matters more than this asymmetry. Even if an American university or dataset creator was not subject to the UK GDPR, that does not determine the lawfulness of my subsequent download, screening, storage and training in London. The CNIL makes the equivalent point for the EU. Each controller’s operation is assessed separately. That is why compliance cannot simply be inherited with a dataset.inheritance2
I should disclose my own interest. This regime makes a provenance-audited corpus a moat. If proven training-data provenance is expensive, anyone who has already paid for it owns an asset that competitors cannot cheaply replicate. I am describing that barrier while building the asset. Provenance is also scientifically useful because it reduces label contamination and permits reproducible source-level ablations. Readers should weigh both facts when judging my argument.
9. What I am asking regulators to clarify
These are the questions I think the ICO could answer now, or Parliament where legislation is needed, in terms useful to every small developer rather than only to me.
1. Transient screening. The ICO should explain whether the reasoning in GC and Others, as assimilated EU case law, extends to genuinely incidental and residual special-category intake in this materially different context, and what preventive and deletion controls are legally required. The CNIL has adopted an analogous approach and draft EDPB Guidelines 03/2026 propose one. Separately, the ICO should identify whether any positive route exists for comparable Article 10 criminal-offence data; the EDPB draft warns that its analysis is not a general exemption and does not clearly supply that route. If existing UK law cannot support an adequately bounded route, Parliament should consider a narrow statutory intake-and-exclusion rule limited to automated classification and exclusion, isolation from training, minimal retention, no ordinary human access, verified deletion, no onward disclosure, and reasonable pre-collection exclusions.
2. Article 11 for unstructured corpora. What a requester must provide; when a provenance vault means the controller is “in a position to identify”; and what a reasonable and proportionate search of an unstructured corpus looks like.
3. Model classes and release contexts. Existing ICO analysis already distinguishes development from deployment, lifecycle stages, and actors, so the gap is narrower than “guidance treats AI as one thing”. What is missing is a sufficiently clear test for whether a low-output discriminative model remains personal data after training. The test should account for model architecture, output type, exposed scores, API versus downloadable release, acceptable privacy-test thresholds and the EDPB’s “insignificant likelihood” logic or a UK equivalent. Guidance should also distinguish source deletion and future-training suppression from the separate, evidence-dependent questions of retraining, unlearning or deleting an existing model when an objection or erasure request succeeds.
4. A worked Article 14 example for large unstructured corpora: what public notice and source disclosure suffice for materially different cohorts, when individual notification is disproportionate, and when it is not. The example should show how a public source registry identifies the source, version and acquisition date with enough specificity to make the notice useful.
5. A UK-interoperable provenance profile building on MLCommons Croissant, SPDX, W3C PROV-O and, where first-party content credentials exist, C2PA. It should carry source, version and acquisition date; licence and rights-evidence snapshots; component lineage; human-authorship or AI-generation evidence; privacy tier; and exclusion history, so that provenance claims become auditable and portable. Although the DSIT call does not focus on copyright, a joint ICO/IPO worked example would reflect how founders actually experience the two rights gates: as one acquisition decision.
Two qualifications matter. A properly conducted DPIA that concludes residual risk is not high is successful compliance. Article 36 prior consultation is mandatory only when high residual risk remains, and I am not asking for pre-clearance of ordinary processing. The ICO Regulatory Sandbox already exists and may provide a useful route for this project. Confidential, project-specific advice, however, cannot replace published guidance that every developer in my position can read and rely on.
10. Am I wrong?
Entirely possible, and in some ways, I hope so. This is a founder’s working analysis, not legal advice, and I would rather be embarrassed in the comments than in an enforcement notice. Corrections are welcome, particularly from data-protection practitioners and anyone at the ICO or DSIT.
I intend to submit a version of the data regulation material in this post to DSIT’s call for evidence on data regulation in the age of AI before it closes on 9 September 2026, with follow-up contact enabled. If you are another founder with a worked example, I encourage you to respond as well. The call explicitly asks for one.
These are good laws that protect things worth protecting, and I am not asking for weaker protection. I am asking for clearer boundaries. The CNIL has begun to answer the hardest question this project raises, and the EDPB has proposed an EU-wide approach. The UK should explain how its existing law applies and consider legislative change only where that answer exposes a genuine gap.
The settlement detail, for those who want it: the final approval order of 20 July 2026 approved a $1.5bn non-reversionary settlement covering 482,460 listed LibGen and PiLiMi works, roughly $3,000 per work before costs and fees, with 440,490 works claimed by 16 April 2026. The order calls it “the largest copyright class action settlement in history”. A work appearing only in Books3, and not on the settlement’s LibGen/PiLiMi Works List, falls outside the settlement class (official class notice). The order records the required destruction of the original LibGen and PiLiMi files and copies originating from them, subject to preservation duties. None of this is a merits judgment pricing acquisition alone; it is the resolution of disputed claims.
A concrete failure mode, described generically: a derivative dataset is finalised on a given date after cleaning and deduplicating an upstream aggregate. Months later, the upstream withdraws one component after rights complaints. Pinned to its finalisation date, the derivative retains whatever survived from that component. A model trained on that derivative may inherit the unresolved risk that surviving records from the component were used, even though the upstream’s current hosting page shows a clean dataset. The withdrawal is a review trigger rather than automatic proof that the older copy or its use is unlawful; that depends on the reason for withdrawal and the licence or permission previously granted. Nothing in that chain identifies which records survived or establishes that any particular actor behaved unlawfully. The engineering lesson is simply that aggregates need versioned, component-level rights review, and that “the dataset page looks fine today” is not evidence about the copy you trained on.

