---
title: "Forward-Deployed Engineering: The Four-Layer Agent Stack"
slug: "forward-deployed-engineering-production-stack-2026"
author: "ibl.ai Engineering"
date: "2026-09-29 11:00:00"
category: "Premium"
topics: "forward-deployed engineering, enterprise AI deployment, agent identity, agent governance, AI agent architecture"
summary: "In every enterprise agent deployment we have worked on, the same four layers appear: an environment provisioner, an evaluation harness, a governed deployment stack, and simulation. This page publishes them as an ownership checklist — which layer you hold, which your vendor holds, and what breaks when the answer is your vendor."
banner: ""
thumbnail: ""
linkedin: |
  Enterprise AI's deployment problem has nothing to do with models.

  In every enterprise agent deployment we have worked on, the same four layers appear:

  1. An environment provisioner — infrastructure-as-code that stands up an isolated tenant with SSO and secrets management, repeatably.
  2. An evaluation harness — automated scoring against your policies and your domain, not a public leaderboard.
  3. A governed deployment stack — every agent with an identity, scoped permissions, and audit from day one.
  4. Simulation — thousands of synthetic conversations before a real customer sees anything.

  September moved two of those four from differentiator toward default. Microsoft put a persistent Copilot agent behind its own governed identity with tenant-level permissions and audit — in private preview, not general availability. NVIDIA open-sourced a runtime containment boundary under Apache 2.0, available today.

  Here is the question that actually decides your deployment, and almost nobody asks it during procurement:

  For each of those four layers — do you own it, or does your vendor?

  Because a provisioner you cannot read is a provisioner you cannot audit. An evaluation harness inside someone else's platform cannot tell you their model got worse. And a governance layer you rent ends the day the contract does.

  On ibl.ai you own all the code and the data across all four layers. The provisioner, the evals, the agent runtime and the sandbox ship as source you run on your own infrastructure.

  #iblai #AgenticAI #EnterpriseAI #ForwardDeployedEngineering #AIGovernance #AgentOps
---

## The Short Answer

**Forward-deployed engineering means shipping enterprise AI as an engineering engagement rather than a product handoff, and in practice it is four layers: environment provisioning, evaluation, governed agent deployment, and simulation. On ibl.ai you own all the code and the data for every one of those layers, so the provisioner, the evals and the agent runtime are auditable source running on your own infrastructure rather than a vendor's black box.**

The gap between an AI demo and an AI deployment is not model quality. It is that a demo needs one environment to work once, and a deployment needs an environment that can be rebuilt, evaluated, governed and tested before anyone depends on it.

This page does something the other write-ups on forward-deployed engineering do not: it treats the four layers as an **ownership checklist**. For each one, the question is who holds it — and what specifically breaks when the answer is your vendor.

## What does forward-deployed engineering mean for enterprise AI?

It means the vendor's engineers are inside the deployment, not on the other side of a support queue. The work of getting an agent into production is treated as engineering to be done jointly, rather than configuration the customer is expected to finish alone.

The term is borrowed from defense and intelligence software, where the person who wrote the system sits with the people using it. It arrived in enterprise AI because the alternative kept failing in a specific, repeatable way.

An AI platform sold as a product assumes the customer's environment resembles the one the product was built in.

In a regulated enterprise it does not — the identity provider is different, the data cannot leave a boundary, and the approval path for a new outbound network connection is measured in weeks.

We have written about the government version of this failure in [Why Government AI Pilots Succeed and Deployments Don't](/blog/forward-deployed-engineering-government-ai-adoption), and about the market shape in [Forward-Deployed AI: Why Enterprise Agent Success Depends on Engineers in the Room](/blog/forward-deployed-ai-enterprise-agent-success). This page is about the stack itself.

## What are the four layers of a production AI agent stack?

Four, and they appear in roughly this order because each one is a prerequisite for the next:

**1. An environment provisioner.** Infrastructure-as-code — Terraform, Helm — that stands up an isolated environment for a customer or a business unit, with SSO wired and secrets management in place, from a single command.

The test is not whether it exists but whether it is *repeatable*. A provisioner used once is a deployment script. A provisioner used every time is the thing that lets you rebuild the environment after an incident, or stand up a staging copy that actually resembles production.

**2. An evaluation harness.** Automated scoring of agent behavior against scenarios from your own domain, run on every change, before anything reaches a user.

The scenarios are the asset here, not the harness. Public benchmarks measure someone else's tasks; a bank's evaluation set encodes its own policies, its own tools and its customers' own phrasing.

**3. A governed deployment stack.** Containerized agents that carry an identity, scoped permissions and observability from the first deployment rather than added after an audit finding.

**4. Simulation.** Synthetic conversations, at volume, against mocked tools — so multi-step agent behavior can be exercised thousands of times without touching a production backend or a real customer.

Nubank published the clearest public example of layer four: **more than 16,000 simulated conversations** used to screen open-weight model configurations before one went live, in [Screen Before You Serve](https://arxiv.org/abs/2609.30137).

We covered what that paper actually shows in [Nubank Screened 16,000 Simulated Chats Before Going Live](/blog/nubank-simulated-conversations-agent-testing-before-production).

## Why did agent identity become a deployment requirement in 2026?

Because the platforms stopped treating it as optional.

On **25 September 2026** Microsoft showed a rebuilt Copilot organized around three surfaces — Home, Code and Autopilot — in which an Autopilot agent runs under its own governed identity rather than a shared service account, with its own memory, compute and workspace inside the tenant.

Microsoft notes that Autopilot was previously called Scout, so the identity predates the September launch.

The significance is not the feature.

It is the concession the feature encodes: an agent that acts inside your systems has to be an *entity in your directory*, because that is the only way its actions can be traced to a known actor and its permissions scoped like anything else you administer.

Note the shipping status, because the coverage blurs it: Autopilot entered **private preview**, not general availability.

Its governance surface is documented separately, under Agent 365: a central agent registry, access control through Entra, and Purview information protection and DLP, administered at the tenant level rather than by the individual user.

Agent 365 itself has been generally available since 1 May 2026.

Scale is what makes it consequential rather than interesting. Microsoft reported **over 30 million paid Microsoft 365 Copilot seats** in the quarter ended 30 June 2026, in its [Q4 FY26 results](https://www.sec.gov/Archives/edgar/data/0000789019/000119312526323632/msft-ex99_1.htm).

A governance model shipped to that installed base becomes the default expectation for every other vendor's agents, including ours.

## Why is agent containment now infrastructure rather than a prompt?

Because the boundary moved below the model. NVIDIA's Open Agent Safety Platform pairs **OpenShell**, an Apache 2.0 runtime boundary available today, with Sentry, an out-of-band watchdog NVIDIA describes as a reference system design rather than a shipping product.

The direction of travel is what matters for layer three. Containment that lives in a system prompt is advice to a model; containment that lives in the runtime is a property of the environment, enforced whether or not the model cooperates.

We looked at the trade in detail — including the part of that announcement you cannot buy yet — in [Agent Containment Moved Into Silicon. What You Still Own.](/blog/nvidia-open-agent-safety-platform-containment-in-silicon)

The practical version of this in our own stack is the agent sandbox: an agent's code runs in a Linux virtual machine that starts with **no network at all**, opened only to the hosts an organization allowlists, with API secrets the agent can use but never read.

The district-facing version is [Letting a K-12 AI Agent Run Code Without Letting Data Out](/blog/agent-sandboxes-k12-districts-code-execution).

## Which layers of the agent stack does your vendor own, and which do you?

This is the checklist, and it is the reason to read this page rather than the others. Take each layer and name the owner:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Layer</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">What breaks if your vendor owns it</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Environment provisioner</strong></td>
      <td style="padding:0.75rem;">You cannot rebuild your own environment, audit what it grants, or stand it up in a region or an air-gapped network the vendor does not serve.</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Evaluation harness</strong></td>
      <td style="padding:0.75rem;">Your evidence that the agent behaves lives inside the system being evaluated. You cannot independently detect that a model update made it worse.</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Governed deployment</strong></td>
      <td style="padding:0.75rem;">Identity, permissions and audit end when the contract ends, and your compliance posture is only as portable as your vendor's export format.</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>Simulation</strong></td>
      <td style="padding:0.75rem;">You can only screen the models that vendor offers. The experiment Nubank ran — screening open-weight candidates and switching to the winner — is not a setting you have.</td>
    </tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;"><strong>ibl.ai: all four are yours</strong></td>
      <td style="padding:0.75rem;">Full source under a perpetual license, model-agnostic across any LLM, no per-seat pricing — running in your cloud, on-premise, GovCloud, or fully air-gapped.</td>
    </tr>
  </tbody>
</table>

The last row is the whole argument. On ibl.ai you own all the code and the data, which means all four layers are artifacts you hold rather than services you rent.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

## How do you tell a forward-deployed engagement from a support contract?

Ask who writes the provisioner. In a real forward-deployed engagement the vendor's engineers produce infrastructure-as-code that runs in *your* account, against *your* identity provider, and you keep it.

Two further tests separate the two. First, whether the evaluation scenarios are yours to keep and extend — if they live only in the vendor's console, the vendor owns your evidence.

Second, whether you can deploy the result somewhere the vendor does not operate, which is the only real proof the stack is portable.

A support contract is measured in response times. A forward-deployed engagement is measured in artifacts that remain useful after it ends.

## Want your four layers provisioned on infrastructure you own?

We deploy all four — provisioner, evals, governed agent runtime and simulation — as source code you keep. [Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact) — ibl.ai is family-owned and operated from New York, NY.

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
