---
title: "How to Write a Statement of Work for AI Infrastructure"
slug: "how-to-write-a-statement-of-work-for-ai-infrastructure"
author: "ibl.ai Engineering"
date: "2026-08-19 13:00:00"
category: "Premium"
topics: "statement of work, ai procurement, ai contract, acceptance criteria, evaluation set, data rights, ai implementation"
summary: "Most AI statements of work define done as a feature list, which makes acceptance a negotiation. Define it as a held-out evaluation set with a passing threshold, and settle source-code rights in the SOW itself."
banner: ""
thumbnail: ""
linkedin: |
  Most AI statements of work fail on the same two clauses.

  1. They define "done" as a feature list.

  A feature list can be delivered while the system doesn't work. "Implement retrieval over the document corpus" — done, technically, at 40% accuracy. Nobody can say whether that's acceptance or breach, so it becomes a negotiation, and the party with more information wins it.

  Define done as a held-out evaluation set drawn from your own data, with a passing threshold agreed before work starts. Then acceptance is a measurement.

  2. They leave source-code and data rights to the master agreement.

  Funding development does not confer the right to operate or modify what was produced. Organizations discover this at closeout, when they want to change something and find they can't.

  A third one worth adding: name the integration endpoints explicitly. "Integrate with existing systems" is not scope — it's an argument scheduled for month four.

  The deeper point is that a good SOW mostly protects you from uncertainty about work nobody has done yet. Reduce what's unbuilt and there's less to protect against.

  On ibl.ai the base already exists — you own all the code and the data, run it model-agnostic across any LLM, with no per-seat pricing — so the SOW scopes integration against named systems instead of construction.

  #iblai #AgenticAI #EnterpriseAI #Procurement #AI #CIO
---

## The Short Answer

**An AI statement of work should define done as a held-out evaluation set with a passing threshold, name the integration endpoints explicitly, and settle source-code and data rights in the SOW itself. On ibl.ai those clauses are simpler because the base already exists — you own all the code and the data, run it model-agnostic across any LLM, with no per-seat pricing, and can deploy anywhere.**

Statements of work for AI infrastructure fail in a small number of predictable ways, and the failures are visible in the document before the engagement starts.

They specify outputs that depend on data quality nobody has assessed. They leave acceptance undefined, so nothing can be closed. They describe integration in the abstract. And they defer ownership to a master agreement that turns out not to say what everyone assumed.

Each is fixable in a paragraph. Here is what those paragraphs need to say.

## How should an AI statement of work define "done"?

As a measurement, not a list of features.

A feature list can be fully delivered while the system does not work. "Implement retrieval over the document corpus" is satisfied by retrieval that returns the right passage 40% of the time.

Nothing in the sentence says otherwise, so acceptance becomes an argument, and the party with more information about the system usually wins it.

The alternative is straightforward. Agree a held-out evaluation set drawn from your own data, with an agreed passing threshold, before work begins. State it in the SOW as the acceptance criterion.

That single change does more than any other clause.

It converts acceptance from negotiation into measurement, gives both parties the same view of progress mid-engagement, and — in public-sector acquisitions — supplies the objective measure usually said to be missing when incentive or fixed-price structures are dismissed.

It also has to be built from real data. An evaluation set assembled by the vendor from convenient examples measures the wrong thing.

## What has to be named rather than described?

The integration endpoints, individually.

"Integrate with existing systems" is not scope.

It is an argument scheduled for month four, when it emerges that the SIS integration everyone assumed was read-only needs to write back, or that the identity provider does not expose the group memberships the permissions model requires.

Name the systems: the SIS, EHR, CRM, ticketing system, identity provider, document stores. Name the direction of data flow for each. Name who owns access to each one, because obtaining credentials is a schedule risk that lands on the buyer more often than the vendor.

This is also the portion of AI work that is genuinely estimable, which is why naming it matters commercially as well as technically. Integration against a defined endpoint list can be scoped and priced. Integration in the abstract cannot.

## Where do source-code and data rights belong?

In the statement of work, negotiated explicitly, not in the master agreement's defaults.

Funding development does not confer the right to operate or modify what was produced. This is the clause organizations most often discover too late — typically at closeout, when they want to change something and find they cannot without returning to the contractor.

Four things need stating. Whether source-code access is required and at what level. Whether your team may modify and redeploy independently. Who owns custom developments built during the engagement.

And what happens to data — inputs, outputs, embeddings, logs — including deletion and who certifies it.

That last one is becoming a formal requirement rather than good practice.

GSA's draft AI clause for federal contracts would require Government Data to be segregated, deleted at contract conclusion, and certified as deleted in writing, which we covered in [GSA's Draft AI Clause and What It Demands Architecturally](/resources/guides/gsa-ai-clause-552-239-7001-data-rights).

Embeddings derived from your data are the store most often forgotten in a deletion clause, and the hardest to reason about afterwards.

## What should the SOW say about models?

That the model layer is replaceable, and who bears the cost when it changes.

Models are deprecated, repriced, and updated on the provider's schedule, and behaviour can shift under a prompt that was working last quarter. A SOW that names a specific model without addressing change has assigned that risk to nobody, which in practice means to you.

State how the system routes across models, what happens when a provider deprecates one, and whether the architecture permits substituting a different model without rewriting the integrations.

If the answer to the last question is no, you have bought a dependency rather than a capability.

## Why do these clauses get easier when the platform already exists?

Because most of what a good SOW protects against is uncertainty about work nobody has done yet.

If the engagement is constructing permissions-aware retrieval, an evaluation harness, guardrails, access control and audit logging, then acceptance criteria, timelines, and rights all have to be negotiated for software that does not exist.

That is why 79% of enterprises reported AI cost overruns in the past twelve months, with 80–85% missing infrastructure forecasts by more than 25%.

When the platform already runs, the SOW scopes integration against a named endpoint list and the agents specific to your organization. The evaluation set is still essential; the timeline is estimable; the rights question is answered by the licence rather than by negotiation.

That is the shape of an ibl.ai engagement: the platform is in production with 1.6M+ users from 400+ organizations and ships with the full source code, so the document is shorter and the parts that usually fail are already settled.

The full contracting picture is in [Time & Materials for AI Infrastructure](/resources/guides/time-and-materials-ai-infrastructure), and the reason hourly billing dominates this market in [Time and Materials Is an Admission, Not a Pricing Model](/blog/time-and-materials-is-an-admission-nobody-can-estimate-the-work).

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
