Generative AI Training Data & Copyright Fair Use Rulings: Where Courts Draw the Line

October 9, 2026

The legal framework surrounding generative artificial intelligence is shifting rapidly from theoretical debate to hard judicial precedent. As technology developers assemble web-scale datasets to train frontier foundation models, copyright owners—ranging from major visual art licensing platforms and book publishers to music labels and software developers—are testing the boundaries of federal copyright protection through high-stakes class-action litigation. 

At the core of virtually every defense raised by model developers stands Section 107 of the U.S. Copyright Act: Fair Use. 

While the defense has historically protected transformative technological tools, federal judges are taking a granular look at how machine learning pipelines ingest, store, and process copyrighted works. For corporate leadership, intellectual property counsel, and commercial AI ventures, understanding how courts weigh dataset ingestion against commercial market harm is essential to mitigating exposure and structuring compliant data pipelines. 

As explored in our ongoing analysis of the shifting battleground in AI copyright law, recent enforcement trends are pushing beyond initial pre-training ingestion into downstream Retrieval-Augmented Generation (RAG) architectures, model fine-tuning, and algorithmic weight retention. 

 

The Four-Factor Fair Use Test in the Context of AI Training 

When evaluating whether scraping and processing copyrighted material for machine learning model training constitutes fair use, federal courts weigh four statutory factors established under 17 U.S.C. § 107, guided by administrative frameworks from the U.S. Copyright Office AI Initiative: 

  1. Purpose and Character of the Use (Factor 1)

When evaluating the four statutory factors under 17 U.S.C. § 107, courts examine whether an AI system creates a genuinely transformative purpose or merely repackages protected expression. Model developers contend that ingesting raw files functions as "intermediate copying"—conceptually akin to search engine web crawlers indexing text to create search parameters. In this view, works are converted into mathematical vector embeddings to extract statistical relationships, syntax rules, and semantic patterns rather than to copy artistic expression. 

However, transformativeness is not an absolute shield. Where AI systems are engineered to generate outputs that mirror the style, structure, or recognizable elements of source material without adding distinct critical or creative commentary, judges are reluctant to grant Factor 1 protection. 

 

  1. Nature of the Copyrighted Work (Factor 2)

This factor looks at the underlying material being scraped. Factual, analytical, historical, and purely technical texts receive a narrower scope of copyright protection than highly expressive creative works, such as fiction novels, fine art photography, commercial illustration, and recorded musical compositions. Mass web-scraping pipelines that indiscriminately collect expressive works lean heavily against a fair use defense under Factor 2. 

  1. Amount and Substantiality Used (Factor 3)

Training competitive large language models (LLMs) or latent diffusion models requires copying entire files—often hundreds of billions of tokens or visual images—to construct pre-training corpora like LAION or Common Crawl. While copying 100% of a work usually weighs against fair use, courts have historically excused full copying in limited technical contexts, such as reverse engineering software to achieve interoperability, provided the end product does not directly compete with the original work. 

  1. Effect on the Potential Market or Value (Factor 4)

Factor 4 is increasingly becoming the primary battleground in generative AI litigation. Courts evaluate whether widespread ingestion destroys an active or potential licensing market for the original works. If content owners can demonstrate that unauthorized scraping undercuts established commercial licensing ecosystems—such as stock photography subscriptions, syndicated journalism feeds, or specialized data APIs—courts view the copying as market substitution, which strongly favors plaintiffs. As administrative policy frameworks monitored by the USPTO AI & Emerging Technologies Hub continue to highlight, evaluating licensing models and market impact remains central to the evolving IP policy debate. 

 

Key Judicial Precedents & Emerging Legal Trends 

Federal courts across multiple jurisdictions are drawing distinct lines between raw data processing and actionable copyright infringement: 

Input Ingestion vs. Substantially Similar Outputs 

Judges consistently separate the mechanics of model training from the generated output. While processing source material to extract statistical weights is frequently recognized as transformative intermediate copying, systems that output verbatim passages, recognizable characters, or near-identical image compositions face significant direct infringement liability. 

Licensing Markets and Commercial Harm 

The commercial viability of the fair use defense closely aligns with market dynamics. In cases like Thomson Reuters v. ROSS Intelligence, courts have scrutinized whether competing platforms used proprietary legal databases to build commercial alternatives, emphasizing that substituting a primary market negates fair use defenses. As coalitions of publishers and copyright owners form structured data-licensing marketplaces, uncompensated web scraping faces heightened judicial skepticism under Factor 4. 

Digital Millennium Copyright Act (DMCA) Section 1202 Claims 

Beyond core copyright infringement, plaintiffs routinely allege violations under Section 1202 of the DMCA. These claims assert that scraping pipelines systematically strip Copyright Management Information (CMI), including author attributions, terms of service, watermarks, and copyright notices, during dataset parsing. While courts have dismissed some CMI claims for lack of explicit intent, developers who intentionally bypass or strip metadata face persistent statutory risk. 

Terms of Service and Website Terms Enforcement 

Parallel to statutory copyright claims, courts are increasingly evaluating breach-of-contract claims stemming from automated web scraping. When developers scrape content behind user authentication walls or in direct violation of explicit website terms of service and robots.txt protocols, courts may enforce contractual restrictions regardless of whether the copying itself qualifies as fair use under copyright law. 

Practical Risk Mitigation for AI Model Developers & Enterprises 

To build legally resilient AI architectures and insulate commercial operations from copyright litigation, organizations should implement rigorous dataset governance practices: 

  1. Maintain Comprehensive Data Lineage Records: Keep complete, time-stamped audit trails for every training dataset. Document source provenance, scraping protocols, filtering criteria, and opt-out compliance to demonstrate good-faith data handling. 
  2. Deploy Strict Inference Guardrails: Implement automated filters on model outputs to detect and suppress verbatim text duplication, copyrighted logos, trademarked assets, and distinct artistic styles prior to user delivery. 
  3. Prioritize Explicit Data Licensing: Reduce reliance on unvetted public web scraping by securing commercial data agreements, utilizing open-access public domain databases, or developing licensed synthetic datasets. 
  4. Respect Robots.txt and Machine-Readable Opt-Outs: Configure automated web crawlers to strictly honor standardized opt-out headers and site governance files to mitigate contract and DMCA exposure. 
  5. Execute Zero-Data-Retention (ZDR) Agreements: When integrating third-party models or processing proprietary client inputs, adopt strict ZDR enterprise agreements to prevent inadvertent data scraping or downstream model training. 

Strategic Guidance for AI Innovators 

Mitigating risks across fair use boundaries, dataset licensing, and generative AI compliance requires a proactive intellectual property strategy. For strategic counsel on dataset governance, AI litigation defense, or technology transactions: