---
title: "NYT v. OpenAI Discovery and the Law Firm AI Checklist"
slug: "nyt-v-openai-unsealed-discovery-law-firm-ai-privilege"
author: "Jaione Amigot"
date: "2026-09-19 12:00:00"
category: "Premium"
topics: "legal AI, attorney-client privilege, NYT v. OpenAI, third-party subpoena, data residency, audit trail, self-hosted AI"
summary: "The unsealed filings reported on September 17, 2026 in NYT v. OpenAI are about training data, not client files. The user-log question was decided earlier on the same docket: 20 million ChatGPT logs, affirmed January 5, 2026."
banner: ""
thumbnail: ""
linkedin: |
  Unsealed New York Times filings against OpenAI, reported September 17, 2026, are about training data, not client files.

  The two get merged into one talking point, and only one is about what a vendor does with your files.

  The unsealed brief concerns how training corpora were assembled: a dataset alleged to hold copies of at least 160,903 unique works from the news plaintiffs.

  The customer-data question was decided earlier on the same docket. On May 13, 2025 a magistrate judge ordered OpenAI to preserve output log data that would otherwise be deleted, including chats users had deleted. ChatGPT Enterprise, Edu and zero-data-retention API were carved out; Free, Plus, Pro, Team and non-ZDR API were not. On November 7, 2025 the court ordered 20 million de-identified logs produced, affirmed January 5, 2026.

  So state the buying criteria precisely, because "zero subpoena surface" is a slogan:

  → Self-hosting removes the vendor as a custodian — no third-party copy to preserve, sample or produce
  → It does not remove the firm from discovery; AI work product on your systems is your ESI
  → A residency clause says where data sits, not who can be compelled to produce it
  → An audit trail is court-grade only if it records who asked, what was retrieved, and which model answered

  With ibl.ai you own all the code and the data — self-hosted inside your own perimeter, model-agnostic across any LLM, usage-based with no per-seat pricing, deployable anywhere from your own cloud to a fully air-gapped network.

  #iblai #AgenticAI #EnterpriseAI #LegalAI #LegalTech #Privilege
---

## The Short Answer

**The unsealed filings reported on September 17, 2026 in NYT v. OpenAI are about training data, not client files. The user-data question was decided earlier on the same docket: 20 million de-identified ChatGPT logs ordered produced, affirmed January 5, 2026. Self-hosting is not "zero subpoena surface" — with ibl.ai you own all the code and the data, so no vendor holds a copy to be subpoenaed.**

A law firm evaluating AI is evaluating a custodian. The litigation record now shows what that costs when the custodian is someone else.

## What did the NYT v. OpenAI filings reported on September 17, 2026 actually show?

An unredacted version of the news plaintiffs' brief in *The New York Times Co. v. Microsoft Corp. and OpenAI*, No. 1:23-cv-11195 (S.D.N.Y.), which had previously been filed under seal.

Its subject is training data acquisition. It quotes Microsoft director Brent Hecht describing AI training on scraped content, in January 2023, as ["the largest theft of labor in human history"](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/).

It alleges that a dataset assembled under an internal initiative called Project Mango contained copies of **at least 160,903 unique works** from the news publishers, and that mid-training datasets held **more than 91,692 copies** of works from the Times, the Daily News and the Center for Investigative Reporting.

Here is the correction worth making up front, because the two are routinely merged into one talking point. That filing is about how a vendor built its corpus from third-party publishers. It is not about what the vendor does with a paying customer's documents.

Both matter to a firm. Only one of them is a discovery story about customer data — and that one is older.

## Did the NYT v. OpenAI orders reach ordinary users' ChatGPT logs?

Yes, on a separate track of the same case, and the sequence is the part worth knowing.

On **May 13, 2025**, Magistrate Judge Ona T. Wang ordered OpenAI to preserve and segregate all output log data that would otherwise be deleted on a going-forward basis — including conversations users had deleted.

OpenAI's COO objected publicly that the requirement conflicted with the privacy commitments the company had made to its users.

The scope was not uniform. [ChatGPT Enterprise, ChatGPT Edu and API endpoints under zero-data-retention agreements were carved out](https://venturebeat.com/ai/sam-altman-calls-for-ai-privilege-as-openai-clarifies-court-order-to-retain-temporary-and-deleted-chatgpt-sessions); Free, Plus, Pro, Team and non-ZDR API traffic were not.

On **October 9, 2025**, Judge Wang [terminated the going-forward obligation](https://www.yahoo.com/news/articles/judge-lifts-order-requiring-openai-151518813.html), with OpenAI no longer required to retain logs past September 26, 2025. Logs already preserved remained accessible, and accounts flagged by the Times still had to be retained.

On **November 7, 2025** the court ordered production of a **20-million-log de-identified sample**, and on **January 5, 2026** Judge Sidney Stein [affirmed it](https://natlawreview.com/article/when-chats-become-evidence-court-affirms-order-requiring-openai-produce-20-million), reasoning that users had voluntarily provided their data as part of ordinary platform usage.

In July 2026 the news plaintiffs went further and [moved for sanctions](https://techcrunch.com/2026/07/09/new-york-times-says-openai-hid-evidence-in-chatgpt-copyright-trial/), alleging that OpenAI had assembled a database of roughly 78 million de-identified conversations before suit was filed. Those are allegations, contested by OpenAI, not findings.

For more than four months, on consumer and standard API tiers, deletion did not mean deletion. The people whose conversations were frozen were not parties to the case and were not in the room when it was argued.

## Does self-hosting AI give a law firm zero subpoena surface?

No, and the slogan should be retired. What self-hosting changes is specific, and worth stating exactly.

**It removes the vendor as a custodian.** If prompts, retrieved documents and model outputs never leave the firm's infrastructure, there is no vendor-side copy to preserve, sample, de-identify or produce. A third-party subpoena or a preservation order directed at the AI provider reaches nothing of the firm's, because the provider holds nothing of the firm's.

**It does not remove the firm from discovery.** AI-generated work product sitting on the firm's own servers is the firm's electronically stored information, subject to the same litigation holds, preservation duties and production obligations as email and the document management system. Running the stack yourself moves the obligation; it does not dissolve it.

**It does not manufacture privilege either.** [ABA Formal Opinion 512](https://www.lawnext.com/wp-content/uploads/2024/07/aba-formal-opinion-512.pdf), issued July 29, 2024, frames the duty as a pre-input analysis: before inputting information relating to a client's representation into a generative AI tool, lawyers must evaluate the risk that it will be disclosed to or accessed by others outside the firm.

The opinion also requires evaluating disclosure inside the firm, and requires the client's informed consent before inputting representation information into a self-learning tool; boilerplate engagement-letter language is not enough. Self-hosting resolves neither.

Self-hosting makes the outside-the-firm half of that evaluation short and answerable. It does not remove the requirement to make it.

The honest version of the claim is arithmetic rather than rhetoric. The number of custodians goes from two to one, and the one that remains is the one already bound by the duty of confidentiality.

## Do data residency guarantees survive a court order in the vendor's jurisdiction?

A residency commitment answers where data is stored. It does not answer who can be compelled to produce it.

The preservation order is the worked example, and nothing about it turned on geography. A U.S. court told a U.S. company to stop deleting records, and its published retention behavior yielded.

The enterprise carve-outs did hold. But they held because a court drew that line in a dispute the customer did not control — not because a contractual commitment executed itself.

So the procurement question for a firm is narrower than the one usually asked. Not "where does this data live," but "who can be ordered to hand it over, and will the firm be a party to the proceeding where that is decided."

## What does a court-grade audit trail for legal AI have to record?

Four things, and a record missing any of them is a usage log rather than evidence.

- **Who asked, and under what authorization.** The identity of the user and the role-based permissions in force at the time of the query.
- **What was retrieved.** The specific documents the answer was grounded in, with identifiers and dates — not a similarity score after the fact.
- **Which model answered, and when.** On a [model-agnostic](/product/agentic-os) platform the model in use changes, so the trail has to pin the version that produced a given output.
- **What was returned, retained under the firm's own schedule.** Retention set by the firm's records policy, exportable on demand.

The retrieval half is not a compliance nicety.

A court in Washington, D.C. [struck a brief filed for Deutsche Bank on September 3, 2026](/blog/deutsche-bank-law-firm-ai-hallucinated-citations-retrieval) because four cited authorities did not exist, which is what happens when a citation is a string a model produced rather than a document someone can open.

## How does ibl.ai deploy AI inside a law firm's perimeter?

By running the whole platform on the firm's own infrastructure, so there is no second custodian to discover.

With ibl.ai you own all the code and the data.

The firm self-hosts the entire stack with full source code access, runs it model-agnostic across any LLM and switches anytime, pays by usage with no per-seat pricing, and can deploy anywhere — its own cloud, on-premise, GovCloud, or a [fully air-gapped network](/blog/air-gapped-ai-for-law-firms-protecting-privilege) with no outbound connectivity.

Governance is built at the platform layer rather than requested of the model: role-based access control, audit logs, retention, and data residency options, with every query recorded against the firm's own schedule.

For a security committee the property that matters is inspectability. Counsel can read the code that touches privileged material instead of accepting an attestation about it, which is a materially different answer to give a client.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University. See how it is configured for firms at [ibl.ai/solutions/legal](/solutions/legal).

ibl.ai is family-owned and operated from New York, NY.

*Related reading: [air-gapped AI for law firms: protecting privilege](/blog/air-gapped-ai-for-law-firms-protecting-privilege) — the deployment pattern this litigation record argues for; and [a court struck a Deutsche Bank brief over fake AI cites](/blog/deutsche-bank-law-firm-ai-hallucinated-citations-retrieval) — why the audit trail has to record retrieved documents, not just answers.*

*Sources: the September 17, 2026 report of the unsealed filings, the Project Mango figure and the Hecht quotation from [TechCrunch](https://techcrunch.com/2026/09/17/microsoft-exec-called-ai-scraping-the-largest-theft-of-labor-in-human-history-new-unredacted-filings-reveal/), with the case timeline and docket number from [Wikipedia's case summary](https://en.wikipedia.org/wiki/The_New_York_Times_v._Microsoft_and_OpenAI); the May 13, 2025 preservation order and its product carve-outs from [VentureBeat](https://venturebeat.com/ai/sam-altman-calls-for-ai-privilege-as-openai-clarifies-court-order-to-retain-temporary-and-deleted-chatgpt-sessions); the October 9, 2025 termination from [Yahoo News](https://www.yahoo.com/news/articles/judge-lifts-order-requiring-openai-151518813.html); the November 7, 2025 production order and the January 5, 2026 affirmance from [The National Law Review](https://natlawreview.com/article/when-chats-become-evidence-court-affirms-order-requiring-openai-produce-20-million); the July 2026 sanctions allegations from [TechCrunch](https://techcrunch.com/2026/07/09/new-york-times-says-openai-hid-evidence-in-chatgpt-copyright-trial/); and the confidentiality duty from [ABA Formal Opinion 512](https://www.lawnext.com/wp-content/uploads/2024/07/aba-formal-opinion-512.pdf).*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
