The Short Answer
Most enterprise AI pilots fail on infrastructure, not model quality: MIT's Project NANDA found that 95% of generative AI pilots produced no measurable P&L impact. ibl.ai closes that gap with an agentic AI platform where you own all the code and the data β self-hosted inside your own perimeter, model-agnostic across any LLM, and usage-based with no per-seat pricing, so you can deploy anywhere.
The failure is structural, not technical. A pilot built against sanitized data in a sandbox has no path to the systems that hold the value.
The organizations that cross into production start from the connective tissue β identity, data pipelines, guardrails, audit β and treat the model as the swappable part.
Why do 95% of enterprise AI pilots produce no P&L impact?
Research consistently shows that 95% of enterprise AI pilots produce no measurable P&L impact. The instinct is to blame the models β upgrade to the newest frontier release, try a different vendor, run another benchmark.
But the models aren't the problem. They've been good enough for most enterprise tasks for a while now.
The number comes from MIT's Project NANDA report The GenAI Divide: State of AI in Business 2025, which examined an estimated $30β40 billion of enterprise investment. It drew on 52 executive interviews, surveys of 153 leaders, and analysis of 300 public AI deployments.
Worth stating plainly: that report is preliminary and not peer-reviewed, and it has been criticized for a short measurement window. The direction is what holds up, and it is corroborated elsewhere.
Its own framing is the "GenAI Divide" β high adoption, low transformation. More than 80% of organizations have piloted tools like ChatGPT or Copilot and nearly 40% report a deployment, yet the gains land on individual productivity rather than the income statement.
What is the infrastructure gap that kills enterprise AI pilots?
The real failure mode is almost always the same: pilots get built in isolation, tested on sanitized data, and demonstrated in controlled environments.
Then someone asks the obvious question β "can this connect to our CRM, our ERP, our proprietary data warehouse?" β and the answer reveals the infrastructure gap nobody planned for.
The model is the cheapest part of the stack. The expensive part is the connective tissue: authentication, data pipelines, security controls, audit logging, and the governance layer that lets a risk officer actually sign off on deployment.
Deloitte's agentic-readiness survey measures the same gap from the organizational side. Only 5% of organizations say their business processes are highly prepared for AI agents, and just 15% have scaled orchestrated, cross-functional multi-agent adoption.
Meanwhile 74% of leaders expect nearly half of their business processes to be redesigned around AI agents within four years. That distance between 5% ready and 74% expecting is where pilots go to die.
What does a pilot cost versus a production deployment?
Pilot economics and production economics are different shapes, and per-seat licensing is what breaks the transition. A pilot with 50 seats is cheap at any price. The same tool at 5,000 employees is a budget line nobody approved.
Per-seat AI tools bill on headcount, not on use β so cost scales with how many people you employ, regardless of whether they touch the system in a given month.
| Approach | List price | 50-seat pilot | 5,000 employees |
|---|---|---|---|
| ChatGPT Enterprise | ~$60/user/mo | $36,000/yr | $3,600,000/yr |
| Glean | ~$40/user/mo | $24,000/yr | $2,400,000/yr |
| Microsoft Copilot | ~$30/user/mo | $18,000/yr | $1,800,000/yr |
| ibl.ai (self-hosted, usage-based) | Tokens + infrastructure | Tracks actual use | Tracks actual use |
The point is not that one vendor is expensive. It is that per-seat pricing is the wrong shape for a production deployment: the bill grows with headcount while the value grows with usage, and those two curves separate the moment a pilot succeeds.
How can you tell which AI pilot will survive contact with production?
Five questions separate a pilot that scales from one that ends as a slide. Each one is about infrastructure, and none of them is about the model.
Identity. Does the system authenticate against your existing directory and inherit its permissions, or does it maintain a second list of who can see what?
Data path. Can it reach the CRM, ERP, and warehouse through governed connections β or does the demo run on an exported spreadsheet?
Audit. When a regulator asks what the system did on a specific date for a specific record, can you answer from logs you hold?
Portability. If the model provider changes price or deprecates a version, is that a configuration change or a migration project?
Ownership. When the contract ends, what do you still have β the running system, or an invoice history?
A pilot that answers all five in advance is a production system with a small user count. A pilot that answers none is a demonstration, and demonstrations do not appear in the P&L.
Who owns the AI infrastructure once the pilot reaches production?
This is the question that decides whether the investment compounds. With managed AI platforms, the infrastructure you spent a year integrating stays on the vendor's side of the boundary.
With ibl.ai, you own all the code and the data. The full source runs under a perpetual license on your infrastructure β your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network.
It is model-agnostic by design: run Claude, GPT, Gemini, Llama, Command, or your own fine-tune, and switch providers without rewriting the platform. Billing is usage-based against a cap you set, with no per-seat pricing.
More than 1.6M users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY β a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
The 95% figure is not a verdict on artificial intelligence. It is a measurement of how many organizations bought a model when they needed infrastructure.
