The Short Answer
Forward-deployed engineering means shipping enterprise AI as an engineering engagement rather than a product handoff, and in practice it is four layers: environment provisioning, evaluation, governed agent deployment, and simulation. On ibl.ai you own all the code and the data for every one of those layers, so the provisioner, the evals and the agent runtime are auditable source running on your own infrastructure rather than a vendor's black box.
The gap between an AI demo and an AI deployment is not model quality. It is that a demo needs one environment to work once, and a deployment needs an environment that can be rebuilt, evaluated, governed and tested before anyone depends on it.
This page does something the other write-ups on forward-deployed engineering do not: it treats the four layers as an ownership checklist. For each one, the question is who holds it β and what specifically breaks when the answer is your vendor.
What does forward-deployed engineering mean for enterprise AI?
It means the vendor's engineers are inside the deployment, not on the other side of a support queue. The work of getting an agent into production is treated as engineering to be done jointly, rather than configuration the customer is expected to finish alone.
The term is borrowed from defense and intelligence software, where the person who wrote the system sits with the people using it. It arrived in enterprise AI because the alternative kept failing in a specific, repeatable way.
An AI platform sold as a product assumes the customer's environment resembles the one the product was built in.
In a regulated enterprise it does not β the identity provider is different, the data cannot leave a boundary, and the approval path for a new outbound network connection is measured in weeks.
We have written about the government version of this failure in Why Government AI Pilots Succeed and Deployments Don't, and about the market shape in Forward-Deployed AI: Why Enterprise Agent Success Depends on Engineers in the Room. This page is about the stack itself.
What are the four layers of a production AI agent stack?
Four, and they appear in roughly this order because each one is a prerequisite for the next:
1. An environment provisioner. Infrastructure-as-code β Terraform, Helm β that stands up an isolated environment for a customer or a business unit, with SSO wired and secrets management in place, from a single command.
The test is not whether it exists but whether it is repeatable. A provisioner used once is a deployment script. A provisioner used every time is the thing that lets you rebuild the environment after an incident, or stand up a staging copy that actually resembles production.
2. An evaluation harness. Automated scoring of agent behavior against scenarios from your own domain, run on every change, before anything reaches a user.
The scenarios are the asset here, not the harness. Public benchmarks measure someone else's tasks; a bank's evaluation set encodes its own policies, its own tools and its customers' own phrasing.
3. A governed deployment stack. Containerized agents that carry an identity, scoped permissions and observability from the first deployment rather than added after an audit finding.
4. Simulation. Synthetic conversations, at volume, against mocked tools β so multi-step agent behavior can be exercised thousands of times without touching a production backend or a real customer.
Nubank published the clearest public example of layer four: more than 16,000 simulated conversations used to screen open-weight model configurations before one went live, in Screen Before You Serve.
We covered what that paper actually shows in Nubank Screened 16,000 Simulated Chats Before Going Live.
Why did agent identity become a deployment requirement in 2026?
Because the platforms stopped treating it as optional.
On 25 September 2026 Microsoft showed a rebuilt Copilot organized around three surfaces β Home, Code and Autopilot β in which an Autopilot agent runs under its own governed identity rather than a shared service account, with its own memory, compute and workspace inside the tenant.
Microsoft notes that Autopilot was previously called Scout, so the identity predates the September launch.
The significance is not the feature.
It is the concession the feature encodes: an agent that acts inside your systems has to be an entity in your directory, because that is the only way its actions can be traced to a known actor and its permissions scoped like anything else you administer.
Note the shipping status, because the coverage blurs it: Autopilot entered private preview, not general availability.
Its governance surface is documented separately, under Agent 365: a central agent registry, access control through Entra, and Purview information protection and DLP, administered at the tenant level rather than by the individual user.
Agent 365 itself has been generally available since 1 May 2026.
Scale is what makes it consequential rather than interesting. Microsoft reported over 30 million paid Microsoft 365 Copilot seats in the quarter ended 30 June 2026, in its Q4 FY26 results.
A governance model shipped to that installed base becomes the default expectation for every other vendor's agents, including ours.
Why is agent containment now infrastructure rather than a prompt?
Because the boundary moved below the model. NVIDIA's Open Agent Safety Platform pairs OpenShell, an Apache 2.0 runtime boundary available today, with Sentry, an out-of-band watchdog NVIDIA describes as a reference system design rather than a shipping product.
The direction of travel is what matters for layer three. Containment that lives in a system prompt is advice to a model; containment that lives in the runtime is a property of the environment, enforced whether or not the model cooperates.
We looked at the trade in detail β including the part of that announcement you cannot buy yet β in Agent Containment Moved Into Silicon. What You Still Own.
The practical version of this in our own stack is the agent sandbox: an agent's code runs in a Linux virtual machine that starts with no network at all, opened only to the hosts an organization allowlists, with API secrets the agent can use but never read.
The district-facing version is Letting a K-12 AI Agent Run Code Without Letting Data Out.
Which layers of the agent stack does your vendor own, and which do you?
This is the checklist, and it is the reason to read this page rather than the others. Take each layer and name the owner:
| Layer | What breaks if your vendor owns it |
|---|---|
| Environment provisioner | You cannot rebuild your own environment, audit what it grants, or stand it up in a region or an air-gapped network the vendor does not serve. |
| Evaluation harness | Your evidence that the agent behaves lives inside the system being evaluated. You cannot independently detect that a model update made it worse. |
| Governed deployment | Identity, permissions and audit end when the contract ends, and your compliance posture is only as portable as your vendor's export format. |
| Simulation | You can only screen the models that vendor offers. The experiment Nubank ran β screening open-weight candidates and switching to the winner β is not a setting you have. |
| ibl.ai: all four are yours | Full source under a perpetual license, model-agnostic across any LLM, no per-seat pricing β running in your cloud, on-premise, GovCloud, or fully air-gapped. |
The last row is the whole argument. On ibl.ai you own all the code and the data, which means all four layers are artifacts you hold rather than services you rent.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
How do you tell a forward-deployed engagement from a support contract?
Ask who writes the provisioner. In a real forward-deployed engagement the vendor's engineers produce infrastructure-as-code that runs in your account, against your identity provider, and you keep it.
Two further tests separate the two. First, whether the evaluation scenarios are yours to keep and extend β if they live only in the vendor's console, the vendor owns your evidence.
Second, whether you can deploy the result somewhere the vendor does not operate, which is the only real proof the stack is portable.
A support contract is measured in response times. A forward-deployed engagement is measured in artifacts that remain useful after it ends.
Want your four layers provisioned on infrastructure you own?
We deploy all four β provisioner, evals, governed agent runtime and simulation β as source code you keep. Book a 30-minute demo or talk to the ibl.ai team β ibl.ai is family-owned and operated from New York, NY.