Skip to content

Doctrinal-plus-Empirical Analysis of AI Training-Data Litigation in US Courts

The computational assimilation of copyrighted expressions for algorithmic training creates fundamental tensions within United States intellectual property jurisprudence. Integrating qualitative doctrinal analysis of the statutory four-factor fair use standard with systematic empirical tracking of federal docket developments reveals the procedural pathways and substantive thresholds governing infringement liability. The resulting synthesis establishes concrete frameworks for harmonizing machine learning innovation with authorial protection through standardized licensing mechanisms and targeted statutory adjustments.

Goal of work

To analyze the doctrinal application of fair use and the empirical patterns of training-data litigation in United States federal courts.

Methodology

Mixed-methods doctrinal analysis and systematic coding of federal court dockets, judicial opinions, and statutory claims in AI training litigation.

Scientific novelty

First comprehensive empirical and doctrinal assessment mapping cause-of-action survival, fair use factor weight, and docket trajectories in U.S. AI copyright disputes.

Document Preview

Review the formatting and introduction. The full version will refine the structure for the selected document standard.

PhD Dissertation

Degree:
Doctrinal-plus-Empirical Analysis of AI Training-Data Litigation in US Courts

Author:

Group

First M. Last

Advisor:

Dr. First Last

City, 2026

Contents

Introduction
Chapter 1. Doctrinal Foundations of Copyright and Machine Learning Ingestion
1.1 The Statutory Architecture of Exclusive Rights and Computational Reproductions
1.2 The Evolution of Non-Expressive and Intermediate Copying Doctrines
1.3 The Four-Factor Fair Use Framework Applied to Large-Scale Corpus Assembly
1.4 International Divergence: U.S. Fair Use versus Foreign Text-and-Data Mining Exceptions
Chapter 2. Empirical Framework and Doctrinal Coding Methodology for Training-Data Dockets
2.1 Design of the Judicial Docket Corpus and Case Selection Parameters
2.2 Qualitative Doctrinal Taxonomy: Classifying Input Ingestion versus Output Substantial Similarity
2.3 Quantitative Tracking of Motion Practice, Preliminary Injunctions, and Procedural Dispositions
2.4 Methodological Boundaries, Evidentiary Limitations, and Observational Constraints
Chapter 3. Doctrinal Analysis of Transformative Use and Market Substitution in Federal Decisions
3.1 Judicial Treatment of Transformative Purpose in Algorithmic Weight Extraction
3.2 The Fourth Factor Inquiry: Market Harm, Licensing Custom, and Lost Potential Exploitation
3.3 Nature of the Work and Quantitative Proportion in Wholesale Scraping Claims
3.4 Doctrinal Synthesis of Emerging District and Circuit Jurisprudence
Chapter 4. Empirical Patterns in AI Training-Data Litigation Across U.S. Jurisdictions
4.1 Distribution of Claims Across Visual, Literary, Code, and Multi-Modal Media
4.2 Early Dismissal Rates and the Viability of Direct versus Secondary Infringement Claims
4.3 Interlocutory Outcomes, Class Certification Hurdles, and Evidentiary Burdens
4.4 Judicial Cognizance of Technical Architectures and Data Provenance Records
Chapter 5. Intersecting Statutory Frameworks: DMCA, Contractual Preemption, and Right of Publicity
5.1 Section 1202 CMI Removal and Integrity Claims in Training Scraping
5.2 Terms of Service Enforcement, Breach of Contract, and Copyright Preemption
5.3 Ancillary State-Law Tort Claims and Unjust Enrichment in Model Ingestion
5.4 Systemic Friction Between Statutory Protections and Algorithmic Training Workflows
Chapter 6. Normative Harmonization, Statutory Reform, and Institutional Resolution Mechanisms
6.1 Standardized Licensing Frameworks and Collective Rights Management Models
6.2 Transparency Mandates, Model Provenance Registries, and Judicial Audit Standards
6.3 Legislative Solutions Balancing Innovation Incentives and Authorial Remuneration
Conclusion
Bibliography

Introduction

The systematic ingestion of vast expressive corpora for training artificial intelligence models has destabilized traditional doctrines of intellectual property in the United States. Federal courts face an unprecedented wave of copyright litigation examining whether scraping protected works to optimize neural parameters constitutes non-infringing transformative fair use or actionable wholesale reproduction [1]. This institutional confrontation exposes deep friction between established precedents governing computational data extraction and the unique economic realities of generative model deployment across commercial sectors [4].

Judicial resolution of these disputes remains fragmented across federal district courts, creating severe doctrinal uncertainty for both technology developers and content creators. Early rulings grapple with distinguishing non-expressive technical copying from market-substituting generative outputs, while procedural challenges surround evidentiary proof, class certification, and Digital Millennium Copyright Act claims [6]. Without a rigorous dual analysis of statutory precedent and litigation patterns, scholarly understanding of judicial decision-making in computational copyright conflicts remains largely speculative and incomplete [5].

This dissertation provides a comprehensive doctrinal-plus-empirical examination of all major artificial intelligence training-data disputes filed in United States federal courts. By coupling qualitative case-law analysis of the four-factor fair use standard with systematic tracking of docket entries, motion outcomes, and cause-of-action survival rates, the study identifies the core legal mechanisms driving judicial determinations [4], [7]. The research clarifies statutory boundaries, assesses procedural trajectories, and provides evidence-based guidance for judicial modernization, legislative adaptation, and institutional data governance frameworks [8].

2.2 Qualitative Doctrinal Taxonomy: Classifying Input Ingestion versus Output Substantial Similarity

To structure the qualitative analysis of federal dockets, this methodology implements a bifurcated coding taxonomy that systematically isolates computational input ingestion from downstream output generation. Judicial determinations regarding intermediate copying during machine learning development demand a rigorous, granular separation between the ingestion of training corpora and the eventual synthesis of generated outputs. By incorporating analytical distinctions across the generative artificial intelligence development cycle, the empirical framework categorizes docket claims according to their specific locus within algorithmic workflows (SSRN, 2023). This classification prevents the conflation of wholesale intermediate extraction with traditional non-literal substantial similarity inquiries. The doctrinal coding taxonomy specifically indexes motion practice and dispositive pleadings by categorizing claims into direct input reproduction, secondary liability, and output similarity. Recent federal decisions demonstrate that judicial evaluations of fair use increasingly hinge on whether liability attaches to the unauthorized ingestion of protected works or to subsequent market substitution, as highlighted in litigation concerning large language models such as Richard Kadrey v. Meta Platforms and Bartz v. Anthropic (SSRN, 2026). Tracking these doctrinal nodes enables the coding matrix to record how courts apply statutory fair use factors to technical data ingestion while isolating distinct evidentiary burdens associated with provenance records and model weight extraction (SSRN, 2023). Consequently, this qualitative taxonomy provides an empirical baseline for evaluating how judicial reasoning separates non-expressive technical reproduction from actionable expressive market harm across emerging federal dockets.

References

  1. Research on the Copyright Fair Use of Text Data Mining in Generative Artificial Intelligence Training
    Jiayu Guo, Wei Lin, Xuan Liu
    DOI Link
  2. Generative Artificial Intelligence and the Doctrine of Fair Use: A Critical Analysis of Nepal's Copyright Framework
    Eunice Poudel
    DOI Link
  3. The Use of GenAI in Courts: Generative Artificial Intelligence and Legal Decision Making
    Minahil Saleem
    DOI Link
  4. Re-Examining the Contours of Fair Dealing and Fair Use in Copyright Infringement in the Wake of Generative Artificial Intelligence -A Comparative Analysis of UK, US and India
    Arunabha Banerjee
  5. Fair Use of Training Data in Generative Artificial Intelligence
    Weiyi Xia
  6. Generative Artificial Intelligence Generates Images: Copyright Infringement or Fair Use
    Ganglin Liu
  7. Copyright in Generative AI training: Balancing Fair Use through Standardization and Transparency
    Daniel Rodriguez Maffioli
  8. Judicial Judgment Standard of Generative Artificial Intelligence and Training Data from the Perspective of Fair Use
    Zhenyi Gao

Bibliography

Verified SourcesFormatting StandardsHigh UniquenessPro Models
Launch Offer -25%

Dissertation

APA 7th Edition (Publication Manual)

$26$34
  • 120+ pages
  • High originality drafting
  • Export to Word
  • Correct formatting
  • Public Preview
    A preview by another author cannot be made private. Your work will be private and completely unique.
  • Bibliography (150+, APA 7th Edition)
    +$1
  • Add alternative sources (News, .gov, .edu)

Dissertation

APA 7th Edition (Publication Manual)