2.2 Qualitative Doctrinal Taxonomy: Classifying Input Ingestion versus Output Substantial Similarity
To structure the qualitative analysis of federal dockets, this methodology implements a bifurcated coding taxonomy that systematically isolates computational input ingestion from downstream output generation. Judicial determinations regarding intermediate copying during machine learning development demand a rigorous, granular separation between the ingestion of training corpora and the eventual synthesis of generated outputs. By incorporating analytical distinctions across the generative artificial intelligence development cycle, the empirical framework categorizes docket claims according to their specific locus within algorithmic workflows (SSRN, 2023). This classification prevents the conflation of wholesale intermediate extraction with traditional non-literal substantial similarity inquiries. The doctrinal coding taxonomy specifically indexes motion practice and dispositive pleadings by categorizing claims into direct input reproduction, secondary liability, and output similarity. Recent federal decisions demonstrate that judicial evaluations of fair use increasingly hinge on whether liability attaches to the unauthorized ingestion of protected works or to subsequent market substitution, as highlighted in litigation concerning large language models such as Richard Kadrey v. Meta Platforms and Bartz v. Anthropic (SSRN, 2026). Tracking these doctrinal nodes enables the coding matrix to record how courts apply statutory fair use factors to technical data ingestion while isolating distinct evidentiary burdens associated with provenance records and model weight extraction (SSRN, 2023). Consequently, this qualitative taxonomy provides an empirical baseline for evaluating how judicial reasoning separates non-expressive technical reproduction from actionable expressive market harm across emerging federal dockets.