Statutory Fair Use and the Transformative Ingestion Doctrine
The theoretical classification of automated dataset ingestion within United States copyright jurisprudence turns on the interaction between transformative purpose and commercial market harm. Legal scholars diverge substantially on how the four-factor fair use balancing test applies to machine learning models. Lee (2023) conceptualizes model ingestion as an extension of non-expressive intermediate copying, arguing that extracting statistical correlations from expressive works fulfills a transformative computational function distinct from the underlying expressive content. In contrast, the Knowing Machines Research Project (2023) emphasizes that dataset curation involves extensive pipeline reproduction, curation choices, and systemic data extraction that cannot be decoupled from original expressive labor, thereby challenging simplistic classifications of intermediate processing. Adding to this debate, contemporary analyses of generative systems stress that the sheer scale of unauthorized ingestion alters traditional market dynamics (Fair Use of Training Data in Generative Artificial Intelligence, 2025). When downstream outputs compete directly with creators in primary markets, transformative computational utility is weakened under statutory analysis (Copy, Paste, and Generate: Copyright Law and Fair Use in the Age of Artificial Intelligence, 2025). Consequently, theoretical frameworks in US litigation navigate a doctrinal divide: one approach prioritizes technological innovation through broad intermediate copying exemptions, whereas the competing paradigm demands stricter market accountability to protect original expressive works from automated substitution.