---
title: "Two Decision Models in Nine Days. One You Can Own."
slug: "decision-layer-models-nine-days-own-the-router"
author: "ibl.ai Engineering"
date: "2026-09-28 14:00:00"
category: "Premium"
topics: "decision models, agent routing, enterprise AI architecture, open weights, inference cost, AI agents"
summary: "TypeSafe shipped Jev on 15 September 2026 and Fastino shipped GLiNER2.5-Decide on 24 September. Two vendors, nine days, the same architectural claim: the routing and classification an agent does all day should not run on a frontier model. One of the two is Apache 2.0."
banner: ""
thumbnail: ""
linkedin: |
  Nine days, two vendors, one architecture.

  TypeSafe shipped Jev on 15 September 2026. Fastino shipped GLiNER2.5-Decide on 24 September. Neither model writes prose. Both exist to make the small, constant judgment calls an agent makes all day — route this, classify that, is this allowed, should we escalate — which have been running on frontier models because that was the only thing in the stack.

  When one vendor says the decision layer should be separate, it is positioning. When two unrelated vendors ship it inside nine days, it is an architecture.

  The interesting difference is not the benchmark. It is the licence.

  GLiNER2.5-Decide is 340M parameters under Apache 2.0, and Fastino documents it running on CPUs and in air-gapped environments — 167.3 ms p50 on a 48-vCPU Xeon, 38.3 ms on a V100.

  Worth saying plainly, because the numbers are already being read the wrong way: the two have never been benchmarked against each other. Fastino's suite is its own, and the 57.5% runner-up in it is JevK5 — an independent Apache-2.0 reproduction by a third party that states on its own repo that it is not affiliated with TypeSafe or Jev.

  So there is no scoreboard here. What there is, is a licence: for the first time the layer that decides what your agents do is something you can hold rather than call. A router you rent is a router someone else can reprice, deprecate, or read.

  With ibl.ai you own all the code and the data — and the platform is model-agnostic, so the decision layer is a component you choose, not one you inherit.

  #iblai #AgenticAI #EnterpriseAI #OpenWeights #AIArchitecture
---

## The Short Answer

**Two decision-layer models shipped nine days apart: TypeSafe Jev on 15 September and Fastino GLiNER2.5-Decide on 24 September 2026. The split from generation is a pattern now. GLiNER2.5-Decide is Apache 2.0, and on ibl.ai you own all the code and the data — the router included.**

Most enterprise agent stacks still send every task to one frontier model. Classify the ticket, check the permission, pick the knowledge base, judge the confidence, decide whether to escalate — then, finally, write something. Five of those six steps produce no prose at all.

## What is a decision model, and why is it not just a smaller LLM?

A decision model returns a structured answer rather than text. There is no decoding step, because there is nothing to decode.

Given a passage and a set of typed questions, GLiNER2.5-Decide returns valid answers with probabilities, confidence scores and constraint-feasibility metadata. It can extract spans and relations and enforce rules across related outputs.

It cannot write you a paragraph, and it is not trying to.

That is the architectural point. Routing, triage, tool selection and guardrail checks are the frequent judgment calls inside an agent pipeline, and they have been running on generation models for the same reason everything else did: that was the only thing in the stack.

## What did each vendor actually ship?

Two models with the same thesis and very different distribution.

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;"></th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">TypeSafe Jev</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Fastino GLiNER2.5-Decide</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">Announced</td><td style="padding:0.75rem;">15 September 2026</td><td style="padding:0.75rem;">24 September 2026</td></tr>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">Size</td><td style="padding:0.75rem;">Not published</td><td style="padding:0.75rem;">340M parameters</td></tr>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">How you get it</td><td style="padding:0.75rem;">API, early access — $0.042 per million input tokens, output unmetered</td><td style="padding:0.75rem;">Open weights, downloadable</td></tr>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;"><strong>Licence</strong></td><td style="padding:0.75rem;">Commercial service</td><td style="padding:0.75rem;"><strong>Apache 2.0</strong></td></tr>
    <tr style="border-bottom:1px solid #e5e7eb;"><td style="padding:0.75rem;">Runs air-gapped</td><td style="padding:0.75rem;">No</td><td style="padding:0.75rem;">Yes, documented</td></tr>
  </tbody>
</table>

When one vendor argues the decision layer should be separate, that is positioning. When two unrelated vendors ship it inside nine days, it is an architecture.

## How fast is it really, and on what hardware?

Fast — but the number depends entirely on the machine, and this is where the figure circulating online goes wrong.

Fastino publishes p50 end-to-end latency on short documents across a range of hardware: **38.3 ms on an NVIDIA V100**, 43.4 ms on an L4, 43.6 ms on a T4, 47.3 ms on an A100, and **167.3 ms on a 48-vCPU Intel Xeon Platinum 8581C**.

At 1,024 tokens the same measurements rise to 52.6 ms on an A100 and 75.6 ms on a V100.

Both halves of that table matter, and a summary that takes the fast number and the CPU claim together gets the model wrong. The sub-40 ms figures are **GPU** measurements; on the 48-vCPU Xeon it is **167.3 ms**, roughly 4.4× slower than the V100.

That the model runs on a CPU at all is the genuinely interesting claim. It just does not run at GPU latency there.

## Has anyone benchmarked the two against each other?

No — and the number being passed around suggests otherwise, so it is worth stating clearly.

Fastino reports GLiNER2.5-Decide at **60.1%** average across 17 datasets spanning classification, routing, triage and content understanding, against **JevK5 at 57.5%**, SemIf at 56.4%, GLiFormer at 49.0% and Laya at 46.6%.

**JevK5 is not TypeSafe's Jev.** It is an independent open-weight reproduction of the idea, built by a third party on Qwen3.5 and released under Apache 2.0, and its own repository states that it is [not affiliated with TypeSafe AI or Jev](https://github.com/allebee/jevk5). TypeSafe's model does not appear in these results at all.

Two further caveats, both from Fastino's own write-up: the suite is **its own internal benchmark**, not an independent one, and the baselines are its own selection.

So this is a vendor scoring itself against alternatives it chose — useful as a sanity check that a 340M model is competitive at this task, and not a basis for picking between the two models this post is about.

## So what actually separates them?

The licence, and it is not close.

GLiNER2.5-Decide is Apache 2.0, and Fastino documents it running locally on CPUs and in air-gapped environments. That makes the decision layer a component you hold rather than a service you call — the first time that has been true for this part of the stack.

The consequence is not philosophical. A router you rent can be repriced, deprecated or version-bumped underneath a workflow you have already validated.

It also sees every routing decision your organization makes, which for a regulated buyer is a record of internal operations leaving the building even when no customer data does.

## What does this change about what an enterprise should build?

Stop budgeting as though one model does everything, and keep the decision layer replaceable.

An agent workflow makes many structured judgments per generated response. Pricing that whole shape at frontier rates is how AI budgets end up dominated by work that never needed a large model.

Splitting the layers is the fix, and it is now a choice between at least two shipped implementations rather than a thing to build yourself.

The part worth protecting is optionality. As of this writing Jev is under a fortnight old and GLiNER2.5-Decide is a few days old; the third will be along shortly.

A platform that lets you swap the decision layer without rewriting the agents around it is worth more than any current benchmark leader.

## Why does ownership decide this one?

Because the router is where your operating logic lives, and a rented router is somebody else's copy of it.

On ibl.ai you own all the code and the data. The platform is model-agnostic, so the decision layer is a component you choose and can change.

That includes an open-weight model running entirely inside your own network, which is exactly what an Apache 2.0 model that runs air-gapped makes possible.

Pricing is usage-based with no per-seat pricing, and you can deploy anywhere: your cloud, your VPC, on-premise, or fully air-gapped.

More than 1.6M users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

*Sources: model, licence, hardware latency table and the internal benchmark from [Fastino's launch post](https://fastino.ai/blog/gliner-2-5-decide-open-weight-decision-model); Jev's pricing and announcement date from [TypeSafe's launch post](https://typesafe.ai/blog/introducing-system-one-models-and-jev); JevK5's independence from [its own repository](https://github.com/allebee/jevk5).*

*Related: [Generation Is Commoditized. Judgment Is the New Frontier](/blog/jev-judge-model-evaluation-not-generation-enterprise) — the first of these two models in depth, and who owns the rubric.*

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
