The Short Answer
Two decision-layer models shipped nine days apart: TypeSafe Jev on 15 September and Fastino GLiNER2.5-Decide on 24 September 2026. The split from generation is a pattern now. GLiNER2.5-Decide is Apache 2.0, and on ibl.ai you own all the code and the data β the router included.
Most enterprise agent stacks still send every task to one frontier model. Classify the ticket, check the permission, pick the knowledge base, judge the confidence, decide whether to escalate β then, finally, write something. Five of those six steps produce no prose at all.
What is a decision model, and why is it not just a smaller LLM?
A decision model returns a structured answer rather than text. There is no decoding step, because there is nothing to decode.
Given a passage and a set of typed questions, GLiNER2.5-Decide returns valid answers with probabilities, confidence scores and constraint-feasibility metadata. It can extract spans and relations and enforce rules across related outputs.
It cannot write you a paragraph, and it is not trying to.
That is the architectural point. Routing, triage, tool selection and guardrail checks are the frequent judgment calls inside an agent pipeline, and they have been running on generation models for the same reason everything else did: that was the only thing in the stack.
What did each vendor actually ship?
Two models with the same thesis and very different distribution.
| TypeSafe Jev | Fastino GLiNER2.5-Decide | |
|---|---|---|
| Announced | 15 September 2026 | 24 September 2026 |
| Size | Not published | 340M parameters |
| How you get it | API, early access β $0.042 per million input tokens, output unmetered | Open weights, downloadable |
| Licence | Commercial service | Apache 2.0 |
| Runs air-gapped | No | Yes, documented |
When one vendor argues the decision layer should be separate, that is positioning. When two unrelated vendors ship it inside nine days, it is an architecture.
How fast is it really, and on what hardware?
Fast β but the number depends entirely on the machine, and this is where the figure circulating online goes wrong.
Fastino publishes p50 end-to-end latency on short documents across a range of hardware: 38.3 ms on an NVIDIA V100, 43.4 ms on an L4, 43.6 ms on a T4, 47.3 ms on an A100, and 167.3 ms on a 48-vCPU Intel Xeon Platinum 8581C.
At 1,024 tokens the same measurements rise to 52.6 ms on an A100 and 75.6 ms on a V100.
Both halves of that table matter, and a summary that takes the fast number and the CPU claim together gets the model wrong. The sub-40 ms figures are GPU measurements; on the 48-vCPU Xeon it is 167.3 ms, roughly 4.4Γ slower than the V100.
That the model runs on a CPU at all is the genuinely interesting claim. It just does not run at GPU latency there.
Has anyone benchmarked the two against each other?
No β and the number being passed around suggests otherwise, so it is worth stating clearly.
Fastino reports GLiNER2.5-Decide at 60.1% average across 17 datasets spanning classification, routing, triage and content understanding, against JevK5 at 57.5%, SemIf at 56.4%, GLiFormer at 49.0% and Laya at 46.6%.
JevK5 is not TypeSafe's Jev. It is an independent open-weight reproduction of the idea, built by a third party on Qwen3.5 and released under Apache 2.0, and its own repository states that it is not affiliated with TypeSafe AI or Jev. TypeSafe's model does not appear in these results at all.
Two further caveats, both from Fastino's own write-up: the suite is its own internal benchmark, not an independent one, and the baselines are its own selection.
So this is a vendor scoring itself against alternatives it chose β useful as a sanity check that a 340M model is competitive at this task, and not a basis for picking between the two models this post is about.
So what actually separates them?
The licence, and it is not close.
GLiNER2.5-Decide is Apache 2.0, and Fastino documents it running locally on CPUs and in air-gapped environments. That makes the decision layer a component you hold rather than a service you call β the first time that has been true for this part of the stack.
The consequence is not philosophical. A router you rent can be repriced, deprecated or version-bumped underneath a workflow you have already validated.
It also sees every routing decision your organization makes, which for a regulated buyer is a record of internal operations leaving the building even when no customer data does.
What does this change about what an enterprise should build?
Stop budgeting as though one model does everything, and keep the decision layer replaceable.
An agent workflow makes many structured judgments per generated response. Pricing that whole shape at frontier rates is how AI budgets end up dominated by work that never needed a large model.
Splitting the layers is the fix, and it is now a choice between at least two shipped implementations rather than a thing to build yourself.
The part worth protecting is optionality. As of this writing Jev is under a fortnight old and GLiNER2.5-Decide is a few days old; the third will be along shortly.
A platform that lets you swap the decision layer without rewriting the agents around it is worth more than any current benchmark leader.
Why does ownership decide this one?
Because the router is where your operating logic lives, and a rented router is somebody else's copy of it.
On ibl.ai you own all the code and the data. The platform is model-agnostic, so the decision layer is a component you choose and can change.
That includes an open-weight model running entirely inside your own network, which is exactly what an Apache 2.0 model that runs air-gapped makes possible.
Pricing is usage-based with no per-seat pricing, and you can deploy anywhere: your cloud, your VPC, on-premise, or fully air-gapped.
More than 1.6M users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
ibl.ai is family-owned and operated from New York, NY β a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
Sources: model, licence, hardware latency table and the internal benchmark from Fastino's launch post; Jev's pricing and announcement date from TypeSafe's launch post; JevK5's independence from its own repository.
Related: Generation Is Commoditized. Judgment Is the New Frontier β the first of these two models in depth, and who owns the rubric.