Unsealed Manhattan Court Records Reveal Knowing Use of Pirated Archives

Summary judgment motions and evidentiary exhibits unsealed between September 18 and 21, 2026, in the Southern District of New York (Authors Guild et al. v. OpenAI Inc., Microsoft Corp., Case No. 1:23-cv-10211) show that OpenAI researchers and senior leadership knowingly collected and processed pirated literature starting as early as 2018. The filings, made widely accessible by the Authors Guild on September 21, 2026, detail internal emails, memorandums, and deposition transcripts documenting the commercial development of foundational language systems.

According to the unsealed records, OpenAI downloaded extensive archives from the shadow library Library Genesis (LibGen) to construct the internal 'Books1' and 'Books2' datasets. These corpora served as central pillars for pretraining GPT-3 and GPT-3.5. Evidence indicates that in April 2019, CEO Sam Altman and then-Research Director Dario Amodei demonstrated an early iteration of GPT-3 to Microsoft co-founder Bill Gates and Microsoft Chief Technology Officer Kevin Scott, explicitly noting the incorporation of the LibGen repository.

The internal documents also document concerns surrounding reputation and public exposure. Communications between Dario Amodei and former OpenAI researcher Sam McCandlish highlighted severe anxiety over headlines reporting that OpenAI built its flagships using copyrighted material retrieved from questionable offshore domains.

Project Clear Dataset Deletions and Internal Labor Displacement Memos

The unsealed records detail internal efforts to mitigate legal exposure through data scrubbing. In June 2022, OpenAI Vice President of Research Bob McGrew documented internally that purging the LibGen datasets would be 'very valuable for legal reasons.' This directive materialized into an internal operational initiative designated 'Project Clear,' which carried out the deletion of downloaded Library Genesis archives across internal repositories.

Beyond dataset acquisition, the discovery record demonstrates that company strategists directly projected the economic displacement of creative professionals. In May 2020, then-Policy Director Jack Clark issued an internal memo stating: 'Our work on AI and Creativity is going to increasingly lead to us creating systems that substitute for the labor of [] people... The better we do on GPT-X, the more worried genre fiction authors will become about us substituting for them on Amazon... Our work in this area will make people unemployed.'

Deposition records show that in 2022, literary evaluation hire Tarun Gogineni directed models toward generating concluding volumes for George R.R. Martin's fantasy series. Internal documentation shows Gogineni characterizing vocal author grievances and income destruction as 'acceptable economic disruption' necessary to advance synthetic capabilities.

Legal Exposure: The Vulnerability of Fair Use Against Willful Infringement

Plaintiffs assert that the unsealed records substantiate willful statutory infringement, a standard that severely complicates OpenAI's fair use defense. Under United States copyright statutes, findings of willful infringement allow courts to award statutory damages of up to $150,000 per registered work. Given the hundreds of thousands of copyrighted titles represented within Library Genesis archives, potential cumulative damages pose extraordinary balance-sheet liabilities.

Whether the presiding judge will invalidate the fair use defense based on the illicit commercial acquisition of pirate corpora remains legally undecided. Cross-motions for partial summary judgment are actively pending, with formal hearings and oral arguments scheduled into early 2027.

Legal observers trace close parallels to pre-trial determinations in Bartz v. Anthropic, where judicial skepticism regarding pirated source acquisition contributed to a landmark $1.5 billion settlement. Consequently, the Authors Guild litigation represents a pivotal stress test on whether commercial AI model training can retain fair use protections when pretraining pipelines rely on systematically pirated data.

Developer and Practitioner Reactions: Automation Inevitability Versus Executive Ethics

Practitioner discussions across the technical ecosystem have bifurcated sharply around historical economics and corporate integrity. One segment of developers characterized labor displacement as an unavoidable byproduct of technological progression, arguing that software tools inherently disrupt manual industries in the same manner that mechanical calculators displaced human computers and automobiles replaced horse-drawn carriages.

Conversely, a broad cohort of software engineers and machine learning practitioners expressed deep cynicism toward lab leadership. Critics emphasized the stark divergence between corporate marketing surrounding responsible AI development and internal execution, noting that clandestine initiatives like 'Project Clear' were organized primarily to eliminate evidence of questionable copyright compliance ahead of public scrutiny.

The consensus among technical observers suggests that leadership's internal communications reveal an operational posture driven by public relations containment rather than principled copyright governance, leaving enterprise developers increasingly wary of licensing terms provided by proprietary model vendors.

Strategic and Regulatory Implications for Thai Enterprises

For enterprise executives and Chief Technology Officers in Thailand, the unsealed filings underscore critical vulnerabilities in foundational model supply chains. While the proceedings sit within United States jurisdiction, Thai commercial entities heavily integrate these underlying architectures across enterprise workflows, including customer relationship management, localization pipelines, and automated content infrastructure.

Judicial penalties or expensive statutory settlements could inevitably translate into escalating API unit economics as vendors attempt to offset licensing and compliance liabilities. Furthermore, Thai conglomerates engaging in cross-border commerce face tightening ESG and governance mandates, requiring demonstrable provenance over algorithmic outputs and training corpora.

Corporate legal counsels and enterprise architects in Thailand must scrutinize vendor indemnification protections, auditing whether existing contracts cover downstream copyright liability originating from pretraining data. Simultaneously, technical roadmaps should incorporate model diversity—evaluating transparent open-weight architectures alongside proprietary solutions—to mitigate dependency on litigation-exposed model families.

Why it matters

Evidence of intentional infringement undermines fair use defenses, driving up statutory liability, licensing costs, and data governance compliance mandates for enterprises worldwide building on proprietary foundation models.

Primary material