The Short Answer
Vals AI raised a $40M Series A at a $400M valuation in August 2026 for a product that validates AI systems rather than building them β evidence that governance, not capability, is now the binding constraint on enterprise AI. The harder question is why observability must be reconstructed from outside at all. On ibl.ai you own all the code and the data, so the audit trail is yours by construction.
A category forms when enough buyers share a problem to support a vendor solving only that problem. Governance just cleared that bar.
What does a $400M valuation for an AI evaluation company signal?
That measurement has become a distinct purchase. Vals AI announced a $40 million Series A led by Andreessen Horowitz on 13 August 2026, at a $400 million valuation, with participation from 8VC, Pear VC, Bloomberg, Hudson River Trading, and NextLadder Ventures.
The company builds independent, domain-specific benchmarks that test models on real professional work. Its evaluations have been cited in model cards from OpenAI, Anthropic, Google, Meta, and xAI.
The finding that made the round legible: frontier models failed roughly half of real finance-analyst tasks in the company's own benchmarking. Capability on public leaderboards did not transfer to the job.
That is the whole thesis in one data point. Buyers no longer doubt that models are impressive; they doubt that impressive translates into their workflow, and they will pay a third party to tell them.
Why do 95% of enterprise AI pilots produce no measurable P&L impact?
This figure gets attributed to several analyst firms. Its actual source is MIT Media Lab's Project NANDA, in "The GenAI Divide: State of AI in Business" β built from 52 executive interviews, 153 survey responses, and analysis of more than 300 public AI deployments.
The study found 95% of pilots delivered no measurable P&L impact, with only about 5% of integrated systems creating significant value.
The reported cause matters more than the headline. Failures traced to data foundations, operating-model gaps, and tools that never entered the workflow they were bought to change β not to weak models.
That distinction is what creates a governance market rather than a better-model market. If the blocker were capability, the fix would be waiting for the next release. Because the blocker is integration and measurement, the fix is infrastructure.
What made 2026 the tipping point for AI governance?
Three forces converged, and each one independently increases the amount of AI an enterprise must account for.
Models commoditized. Alibaba's Qwen family crossed 3 billion downloads in six months, passing Meta and Google to become the most-downloaded open model family. Google recorded roughly 418 million downloads over 2026 and Meta 227 million. When frontier-class capability is near-free, the model stops being the differentiator.
Agents proliferated. Goldman Sachs deployed Cognition's Devin across its engineering organization, running hundreds of AI agents alongside roughly 12,000 human engineers. Agents that take actions create obligations a chatbot never did.
Regulation converged. The EU's Cloud and AI Development Act sets sovereignty requirements for public-sector workloads. Kenya's draft National AI Policy proposes spreading liability across developers, deployers, and operators rather than resting it on one party β a structure worth watching precisely because it is still open for comment rather than settled law.
Who governs the AI agents an enterprise has already deployed?
Usually nobody, and that is the honest state of most programs. The question has moved from "should we deploy agents" to "who is accountable for the ones already running."
Shadow AI is the specific version of this. Low-code platforms let HR, finance, and operations teams build agents without going through IT, so the inventory problem precedes the governance problem β you cannot govern a system nobody has enumerated.
An agent that reads and replies is a search interface with better manners. An agent that files, transacts, provisions, or communicates on the organization's behalf is an actor whose actions the organization owns.
Regulators, auditors, and courts will not accept "the vendor's model did it" as a division of responsibility. The obligation lands on the deploying organization regardless of where the inference ran.
Can you govern AI you don't control?
Partially β and the gap is where this category's economics come from.
Four questions sit at the centre of any serious AI audit, and each has an answer determined by architecture rather than by policy:
| Audit question | On a hosted API | Self-hosted and owned |
|---|---|---|
| What did the agent do? | Vendor's log retention | Your logs, your retention |
| Which model version produced it? | May be replaced silently | Pinned until you change it |
| Where did data go at inference? | Contractual assurance | Never left the perimeter |
| What is the blast radius? | Bounded by vendor controls | Bounded by your network |
| Cost at 5,000 users | Per seat, scales with headcount | Usage-based or flat license |
None of this makes independent evaluation redundant. Benchmarking whether a model does the job is genuinely useful whoever hosts it, which is why Vals AI's customers include organizations running their own infrastructure.
The distinction is between evaluation and observability. Evaluation asks whether the system is good enough; you want an outside party for that.
Observability asks what the system actually did; that should not require inference from outside, and on infrastructure you own it does not.
Where ibl.ai fits
ibl.ai is the agentic AI platform where you own all the code and the data.
You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing β so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.
Because the runtime is yours, the governance artifacts are produced rather than reconstructed: tool-call logs, model version pins, and egress boundaries are properties of the deployment.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
Related: AI Governance Platforms: Enterprise Guide 2026 Β· Writing an AI Governance Policy Β· Enterprise AI Agents ROI