ibl.ai Agentic AI Blog

Insights on building and deploying agentic AI systems. Our blog covers AI agent architectures, LLM infrastructure, MCP servers, enterprise deployment strategies, and real-world implementation guides. Whether you are a developer building AI agents, a CTO evaluating agentic platforms, or a technical leader driving AI adoption, you will find practical guidance here.

Topics We Cover

Featured Research and Reports

We analyze key research from leading institutions and labs including Google DeepMind, Anthropic, OpenAI, Meta AI, McKinsey, and the World Economic Forum. Our content includes detailed analysis of reports on AI agents, foundation models, and enterprise AI strategy.

For Technical Leaders

CTOs, engineering leads, and AI architects turn to our blog for guidance on agent orchestration, model evaluation, infrastructure planning, and building production-ready AI systems. We provide frameworks for responsible AI deployment that balance capability with safety and reliability.

Back to Blog

NYT v. OpenAI Discovery and the Law Firm AI Checklist

Jaione AmigotSeptember 19, 2026
Premium

The unsealed filings reported on September 17, 2026 in NYT v. OpenAI are about training data, not client files. The user-log question was decided earlier on the same docket: 20 million ChatGPT logs, affirmed January 5, 2026.

The Short Answer

The unsealed filings reported on September 17, 2026 in NYT v. OpenAI are about training data, not client files. The user-data question was decided earlier on the same docket: 20 million de-identified ChatGPT logs ordered produced, affirmed January 5, 2026. Self-hosting is not "zero subpoena surface" β€” with ibl.ai you own all the code and the data, so no vendor holds a copy to be subpoenaed.

A law firm evaluating AI is evaluating a custodian. The litigation record now shows what that costs when the custodian is someone else.

What did the NYT v. OpenAI filings reported on September 17, 2026 actually show?

An unredacted version of the news plaintiffs' brief in The New York Times Co. v. Microsoft Corp. and OpenAI, No. 1:23-cv-11195 (S.D.N.Y.), which had previously been filed under seal.

Its subject is training data acquisition. It quotes Microsoft director Brent Hecht describing AI training on scraped content, in January 2023, as "the largest theft of labor in human history".

It alleges that a dataset assembled under an internal initiative called Project Mango contained copies of at least 160,903 unique works from the news publishers, and that mid-training datasets held more than 91,692 copies of works from the Times, the Daily News and the Center for Investigative Reporting.

Here is the correction worth making up front, because the two are routinely merged into one talking point. That filing is about how a vendor built its corpus from third-party publishers. It is not about what the vendor does with a paying customer's documents.

Both matter to a firm. Only one of them is a discovery story about customer data β€” and that one is older.

Did the NYT v. OpenAI orders reach ordinary users' ChatGPT logs?

Yes, on a separate track of the same case, and the sequence is the part worth knowing.

On May 13, 2025, Magistrate Judge Ona T. Wang ordered OpenAI to preserve and segregate all output log data that would otherwise be deleted on a going-forward basis β€” including conversations users had deleted.

OpenAI's COO objected publicly that the requirement conflicted with the privacy commitments the company had made to its users.

The scope was not uniform. ChatGPT Enterprise, ChatGPT Edu and API endpoints under zero-data-retention agreements were carved out; Free, Plus, Pro, Team and non-ZDR API traffic were not.

On October 9, 2025, Judge Wang terminated the going-forward obligation, with OpenAI no longer required to retain logs past September 26, 2025. Logs already preserved remained accessible, and accounts flagged by the Times still had to be retained.

On November 7, 2025 the court ordered production of a 20-million-log de-identified sample, and on January 5, 2026 Judge Sidney Stein affirmed it, reasoning that users had voluntarily provided their data as part of ordinary platform usage.

In July 2026 the news plaintiffs went further and moved for sanctions, alleging that OpenAI had assembled a database of roughly 78 million de-identified conversations before suit was filed. Those are allegations, contested by OpenAI, not findings.

For more than four months, on consumer and standard API tiers, deletion did not mean deletion. The people whose conversations were frozen were not parties to the case and were not in the room when it was argued.

Does self-hosting AI give a law firm zero subpoena surface?

No, and the slogan should be retired. What self-hosting changes is specific, and worth stating exactly.

It removes the vendor as a custodian. If prompts, retrieved documents and model outputs never leave the firm's infrastructure, there is no vendor-side copy to preserve, sample, de-identify or produce. A third-party subpoena or a preservation order directed at the AI provider reaches nothing of the firm's, because the provider holds nothing of the firm's.

It does not remove the firm from discovery. AI-generated work product sitting on the firm's own servers is the firm's electronically stored information, subject to the same litigation holds, preservation duties and production obligations as email and the document management system. Running the stack yourself moves the obligation; it does not dissolve it.

It does not manufacture privilege either. ABA Formal Opinion 512, issued July 29, 2024, frames the duty as a pre-input analysis: before inputting information relating to a client's representation into a generative AI tool, lawyers must evaluate the risk that it will be disclosed to or accessed by others outside the firm.

The opinion also requires evaluating disclosure inside the firm, and requires the client's informed consent before inputting representation information into a self-learning tool; boilerplate engagement-letter language is not enough. Self-hosting resolves neither.

Self-hosting makes the outside-the-firm half of that evaluation short and answerable. It does not remove the requirement to make it.

The honest version of the claim is arithmetic rather than rhetoric. The number of custodians goes from two to one, and the one that remains is the one already bound by the duty of confidentiality.

Do data residency guarantees survive a court order in the vendor's jurisdiction?

A residency commitment answers where data is stored. It does not answer who can be compelled to produce it.

The preservation order is the worked example, and nothing about it turned on geography. A U.S. court told a U.S. company to stop deleting records, and its published retention behavior yielded.

The enterprise carve-outs did hold. But they held because a court drew that line in a dispute the customer did not control β€” not because a contractual commitment executed itself.

So the procurement question for a firm is narrower than the one usually asked. Not "where does this data live," but "who can be ordered to hand it over, and will the firm be a party to the proceeding where that is decided."

Four things, and a record missing any of them is a usage log rather than evidence.

  • Who asked, and under what authorization. The identity of the user and the role-based permissions in force at the time of the query.
  • What was retrieved. The specific documents the answer was grounded in, with identifiers and dates β€” not a similarity score after the fact.
  • Which model answered, and when. On a model-agnostic platform the model in use changes, so the trail has to pin the version that produced a given output.
  • What was returned, retained under the firm's own schedule. Retention set by the firm's records policy, exportable on demand.

The retrieval half is not a compliance nicety.

A court in Washington, D.C. struck a brief filed for Deutsche Bank on September 3, 2026 because four cited authorities did not exist, which is what happens when a citation is a string a model produced rather than a document someone can open.

How does ibl.ai deploy AI inside a law firm's perimeter?

By running the whole platform on the firm's own infrastructure, so there is no second custodian to discover.

With ibl.ai you own all the code and the data.

The firm self-hosts the entire stack with full source code access, runs it model-agnostic across any LLM and switches anytime, pays by usage with no per-seat pricing, and can deploy anywhere β€” its own cloud, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

Governance is built at the platform layer rather than requested of the model: role-based access control, audit logs, retention, and data residency options, with every query recorded against the firm's own schedule.

For a security committee the property that matters is inspectability. Counsel can read the code that touches privileged material instead of accepting an attestation about it, which is a materially different answer to give a client.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University. See how it is configured for firms at ibl.ai/solutions/legal.

ibl.ai is family-owned and operated from New York, NY.

Related reading: air-gapped AI for law firms: protecting privilege β€” the deployment pattern this litigation record argues for; and a court struck a Deutsche Bank brief over fake AI cites β€” why the audit trail has to record retrieved documents, not just answers.

Sources: the September 17, 2026 report of the unsealed filings, the Project Mango figure and the Hecht quotation from TechCrunch, with the case timeline and docket number from Wikipedia's case summary; the May 13, 2025 preservation order and its product carve-outs from VentureBeat; the October 9, 2025 termination from Yahoo News; the November 7, 2025 production order and the January 5, 2026 affirmance from The National Law Review; the July 2026 sanctions allegations from TechCrunch; and the confidentiality duty from ABA Formal Opinion 512.

Why does owning the AI stack matter?

ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing β€” so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.

  • You own all the code and the data

    Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform β€” the stack itself is yours.

  • Model-agnostic

    Run any LLM β€” Claude, GPT, Gemini, Llama, Command, or your own fine-tune β€” and switch providers without rewriting the platform.

  • No per-seat pricing

    Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.

  • Deploy anywhere

    Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY β€” a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

See the ibl.ai AI Operating System in Action

Discover how leading universities and organizations are transforming education with the ibl.ai AI Operating System. Explore real-world implementations from Harvard, MIT, Stanford, and users from 400+ institutions worldwide.

View Case Studies
Work with our team

Pilots, deployment, and full ownership

Most enterprise engagements are one-time, not subscriptions. You integrate ibl.ai with your own data, deploy it on your own infrastructure, and the engineering hours scale with the work β€” so the price tracks the scope, not your headcount.

Start here

Pilot

from $15K

fixed scope Β· fixed timeline

A time-boxed proof of value on your real data β€” not a slide deck.

Best for: Teams that want to see ibl.ai working before committing.

  • Deployed on your infrastructure or our cloud
  • 1–2 production agents wired to a slice of your data
  • One integration (LMS / SIS / SSO / data source)
  • Weekly working sessions with our engineers
  • Pilot fee credits toward a full engagement
Scope a pilot
Most common

Integration & Deployment

$25K – $80K

one-time Β· not a subscription

Full deployment integrated with your data and systems. Engineering hours scale with scope.

Best for: Organizations rolling ibl.ai out across a department, campus, or business unit.

  • Platform deployed in your VPC, on-prem, or air-gapped
  • Integrated with your data + identity (SSO / SAML)
  • Multiple custom agents built to your workflows
  • Engineering hours proportional to scope
  • You own the data Β· run any LLM you choose
Plan a deployment
Full ownership

Codebase Transfer + Custom AI Engineering

Six figures

perpetual license Β· you own the stack

We transfer the full source code. You own and self-host the entire platform β€” outright.

Best for: Government, defense, and enterprises that require perpetual ownership and sovereignty.

  • Complete source-code transfer + perpetual license
  • Dedicated AI engineering team on your roadmap
  • Custom agents, models, and integrations to spec
  • Air-gapped capable Β· zero vendor lock-in
  • Family-owned, New York–based long-term partner
Talk about ownership
You own the code and data Run any LLM β€” Claude, GPT, Gemini, Llama Family-owned & operated from New York, NY