ibl.ai Agentic AI Blog

Insights on building and deploying agentic AI systems. Our blog covers AI agent architectures, LLM infrastructure, MCP servers, enterprise deployment strategies, and real-world implementation guides. Whether you are a developer building AI agents, a CTO evaluating agentic platforms, or a technical leader driving AI adoption, you will find practical guidance here.

Topics We Cover

Featured Research and Reports

We analyze key research from leading institutions and labs including Google DeepMind, Anthropic, OpenAI, Meta AI, McKinsey, and the World Economic Forum. Our content includes detailed analysis of reports on AI agents, foundation models, and enterprise AI strategy.

For Technical Leaders

CTOs, engineering leads, and AI architects turn to our blog for guidance on agent orchestration, model evaluation, infrastructure planning, and building production-ready AI systems. We provide frameworks for responsible AI deployment that balance capability with safety and reliability.

Back to Blog

Two Decision Models in Nine Days. One You Can Own.

ibl.ai EngineeringSeptember 28, 2026
Premium

TypeSafe shipped Jev on 15 September 2026 and Fastino shipped GLiNER2.5-Decide on 24 September. Two vendors, nine days, the same architectural claim: the routing and classification an agent does all day should not run on a frontier model. One of the two is Apache 2.0.

The Short Answer

Two decision-layer models shipped nine days apart: TypeSafe Jev on 15 September and Fastino GLiNER2.5-Decide on 24 September 2026. The split from generation is a pattern now. GLiNER2.5-Decide is Apache 2.0, and on ibl.ai you own all the code and the data β€” the router included.

Most enterprise agent stacks still send every task to one frontier model. Classify the ticket, check the permission, pick the knowledge base, judge the confidence, decide whether to escalate β€” then, finally, write something. Five of those six steps produce no prose at all.

What is a decision model, and why is it not just a smaller LLM?

A decision model returns a structured answer rather than text. There is no decoding step, because there is nothing to decode.

Given a passage and a set of typed questions, GLiNER2.5-Decide returns valid answers with probabilities, confidence scores and constraint-feasibility metadata. It can extract spans and relations and enforce rules across related outputs.

It cannot write you a paragraph, and it is not trying to.

That is the architectural point. Routing, triage, tool selection and guardrail checks are the frequent judgment calls inside an agent pipeline, and they have been running on generation models for the same reason everything else did: that was the only thing in the stack.

What did each vendor actually ship?

Two models with the same thesis and very different distribution.

TypeSafe Jev Fastino GLiNER2.5-Decide
Announced15 September 202624 September 2026
SizeNot published340M parameters
How you get itAPI, early access β€” $0.042 per million input tokens, output unmeteredOpen weights, downloadable
LicenceCommercial serviceApache 2.0
Runs air-gappedNoYes, documented

When one vendor argues the decision layer should be separate, that is positioning. When two unrelated vendors ship it inside nine days, it is an architecture.

How fast is it really, and on what hardware?

Fast β€” but the number depends entirely on the machine, and this is where the figure circulating online goes wrong.

Fastino publishes p50 end-to-end latency on short documents across a range of hardware: 38.3 ms on an NVIDIA V100, 43.4 ms on an L4, 43.6 ms on a T4, 47.3 ms on an A100, and 167.3 ms on a 48-vCPU Intel Xeon Platinum 8581C.

At 1,024 tokens the same measurements rise to 52.6 ms on an A100 and 75.6 ms on a V100.

Both halves of that table matter, and a summary that takes the fast number and the CPU claim together gets the model wrong. The sub-40 ms figures are GPU measurements; on the 48-vCPU Xeon it is 167.3 ms, roughly 4.4Γ— slower than the V100.

That the model runs on a CPU at all is the genuinely interesting claim. It just does not run at GPU latency there.

Has anyone benchmarked the two against each other?

No β€” and the number being passed around suggests otherwise, so it is worth stating clearly.

Fastino reports GLiNER2.5-Decide at 60.1% average across 17 datasets spanning classification, routing, triage and content understanding, against JevK5 at 57.5%, SemIf at 56.4%, GLiFormer at 49.0% and Laya at 46.6%.

JevK5 is not TypeSafe's Jev. It is an independent open-weight reproduction of the idea, built by a third party on Qwen3.5 and released under Apache 2.0, and its own repository states that it is not affiliated with TypeSafe AI or Jev. TypeSafe's model does not appear in these results at all.

Two further caveats, both from Fastino's own write-up: the suite is its own internal benchmark, not an independent one, and the baselines are its own selection.

So this is a vendor scoring itself against alternatives it chose β€” useful as a sanity check that a 340M model is competitive at this task, and not a basis for picking between the two models this post is about.

So what actually separates them?

The licence, and it is not close.

GLiNER2.5-Decide is Apache 2.0, and Fastino documents it running locally on CPUs and in air-gapped environments. That makes the decision layer a component you hold rather than a service you call β€” the first time that has been true for this part of the stack.

The consequence is not philosophical. A router you rent can be repriced, deprecated or version-bumped underneath a workflow you have already validated.

It also sees every routing decision your organization makes, which for a regulated buyer is a record of internal operations leaving the building even when no customer data does.

What does this change about what an enterprise should build?

Stop budgeting as though one model does everything, and keep the decision layer replaceable.

An agent workflow makes many structured judgments per generated response. Pricing that whole shape at frontier rates is how AI budgets end up dominated by work that never needed a large model.

Splitting the layers is the fix, and it is now a choice between at least two shipped implementations rather than a thing to build yourself.

The part worth protecting is optionality. As of this writing Jev is under a fortnight old and GLiNER2.5-Decide is a few days old; the third will be along shortly.

A platform that lets you swap the decision layer without rewriting the agents around it is worth more than any current benchmark leader.

Why does ownership decide this one?

Because the router is where your operating logic lives, and a rented router is somebody else's copy of it.

On ibl.ai you own all the code and the data. The platform is model-agnostic, so the decision layer is a component you choose and can change.

That includes an open-weight model running entirely inside your own network, which is exactly what an Apache 2.0 model that runs air-gapped makes possible.

Pricing is usage-based with no per-seat pricing, and you can deploy anywhere: your cloud, your VPC, on-premise, or fully air-gapped.

More than 1.6M users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY β€” a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

Sources: model, licence, hardware latency table and the internal benchmark from Fastino's launch post; Jev's pricing and announcement date from TypeSafe's launch post; JevK5's independence from its own repository.

Related: Generation Is Commoditized. Judgment Is the New Frontier β€” the first of these two models in depth, and who owns the rubric.

Why does owning the AI stack matter?

ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing β€” so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.

  • You own all the code and the data

    Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform β€” the stack itself is yours.

  • Model-agnostic

    Run any LLM β€” Claude, GPT, Gemini, Llama, Command, or your own fine-tune β€” and switch providers without rewriting the platform.

  • No per-seat pricing

    Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.

  • Deploy anywhere

    Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY β€” a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.

Related Articles

Tencent's 770B Hy4 Is Apache-2.0, and 1.56 TB of Weights

Tencent released Hy4 preview on 28 August 2026 under a genuine Apache License 2.0: 770B total parameters, 49B active, 1M context. The licence is permissive, but the BF16 weights are 1.56 TB and an 8xH100 node cannot load either checkpoint.

Miguel AmigotSeptember 13, 2026

The Open-Weight Price Floor Is Now the Market's Floor

Kimi K3 reached frontier-tier benchmarks at roughly a third of frontier pricing. Meta shipped a 30B Apache-2.0 agentic model that runs on one consumer GPU. Anthropic cut Fable-line cache reads 75%. Open weights are now setting the price of closed models.

ibl.ai EngineeringSeptember 7, 2026

NVIDIA's Open Routing Layer: Why the Model Stopped Being the Moat

NVIDIA shipped an efficient open model and an open routing library on the same day. Together they commoditize the model layer and move the durable advantage to the routing layer β€” which is the one piece you should refuse to rent. What routing saves, what open weights do not buy you, and the three layers worth owning.

ibl.ai EngineeringAugust 12, 2026

SaaS Fragmentation Is the Hidden Cost of Enterprise AI

Enterprises run six or seven per-seat tools that each hold a partial copy of the same customer. That fragmentation, not model capability, is what stalls AI deployment β€” and it carries a per-seat bill that grows with headcount. This post itemizes the fragmentation tax and shows the MCP-based orchestration layer that reads across every system instead of adding another one.

Miguel AmigotJuly 31, 2026

See the ibl.ai AI Operating System in Action

Discover how leading universities and organizations are transforming education with the ibl.ai AI Operating System. Explore real-world implementations from Harvard, MIT, Stanford, and users from 400+ institutions worldwide.

View Case Studies
Work with our team

Pilots, deployment, and full ownership

Most enterprise engagements are one-time, not subscriptions. You integrate ibl.ai with your own data, deploy it on your own infrastructure, and the engineering hours scale with the work β€” so the price tracks the scope, not your headcount.

Start here

Pilot

from $15K

fixed scope Β· fixed timeline

A time-boxed proof of value on your real data β€” not a slide deck.

Best for: Teams that want to see ibl.ai working before committing.

  • Deployed on your infrastructure or our cloud
  • 1–2 production agents wired to a slice of your data
  • One integration (LMS / SIS / SSO / data source)
  • Weekly working sessions with our engineers
  • Pilot fee credits toward a full engagement
Scope a pilot
Most common

Integration & Deployment

$25K – $80K

one-time Β· not a subscription

Full deployment integrated with your data and systems. Engineering hours scale with scope.

Best for: Organizations rolling ibl.ai out across a department, campus, or business unit.

  • Platform deployed in your VPC, on-prem, or air-gapped
  • Integrated with your data + identity (SSO / SAML)
  • Multiple custom agents built to your workflows
  • Engineering hours proportional to scope
  • You own the data Β· run any LLM you choose
Plan a deployment
Full ownership

Codebase Transfer + Custom AI Engineering

Custom quote

perpetual license Β· you own the stack

We transfer the full source code. You own and self-host the entire platform β€” outright.

Best for: Organizations and enterprises that benefit from perpetual ownership and sovereignty.

  • Complete source-code transfer + perpetual license
  • Dedicated AI engineering team on your roadmap
  • Custom agents, models, and integrations to spec
  • Air-gapped capable Β· zero vendor lock-in
  • Family-owned, New York–based long-term partner
Talk about ownership
You own the code and data Run any LLM β€” Claude, GPT, Gemini, Llama Family-owned & operated from New York, NY