---
title: "Bad Theory Labs' Interference Search: Check the Scope"
slug: "bad-theory-labs-btl-4-verifying-model-claims"
author: "ibl.ai Engineering"
date: "2026-09-29 13:00:00"
category: "Premium"
topics: "Bad Theory Labs, Interference Search, BTL-4, open-weight models, model evaluation, model-agnostic architecture"
summary: "Bad Theory Labs published a paper, Apache-2.0 code, trained judge weights, raw results and a log mapping every number to the command that produced it. The figures are real measurements — on Countdown, an arithmetic puzzle with an exact solver. The unsupported step is not the lab's; it is the leap from that benchmark to enterprise agent reasoning."
banner: ""
thumbnail: ""
linkedin: |
  A lab called Bad Theory Labs is being credited with "cracking non-linear AI reasoning." I went to check, expecting to find marketing. I found the opposite.

  Interference Search ships with a paper, Apache-2.0 code, the trained judge weights, every raw result, a research log that includes the failed experiments, and a CLAIMS.md that maps each number in the paper to the file and the command that produced it.

  That is more reproducible than a typical preprint. It is not on arXiv and it is not peer-reviewed — but "unverifiable" is the wrong word for it, and I had to correct my own draft to say so.

  So the numbers are measurements, not claims. Here is the part the retelling drops.

  They are measured on Countdown: give the model a few numbers and a target, ask for an expression using each once. It has an exact solver, which is precisely why it is a good research testbed — every state can be labelled alive or dead without an LLM grading it.

  On that task, reasoning over merged states solved 30 of 30 hard problems where a single line of thought solved 21, in 3.0 sequential steps instead of 23.7.

  A real result. Also a result about arithmetic search with a ground-truth oracle.

  Your agents do not run Countdown. They run your policies, against your tools, where nobody can label a state dead.

  So the question is not "is it real?" It is "what would it cost me to find out whether it helps MY workload?" If your platform is a per-seat subscription to one vendor's models, the answer is that you cannot. If you own the stack, it is an afternoon.

  On ibl.ai you own all the code and the data and the platform is model-agnostic across any LLM.

  Credit where it is due: this lab published its failures. Read the scope, not the multiplier.

  #iblai #AgenticAI #EnterpriseAI #OpenWeights #ModelEvaluation #Reproducibility
---

## The Short Answer

**Bad Theory Labs published Interference Search with a paper, Apache-2.0 code, trained judge weights, raw results and a claims-to-evidence log, so its figures are measurements — measured on Countdown, an arithmetic puzzle with an exact solver. On ibl.ai you own all the code and the data and run the platform model-agnostic across any LLM, so testing whether a new method helps your own workload is an afternoon inside your own perimeter rather than a guess from someone else's benchmark.**

The critique behind the work is real. Language models reason in a line, and when a step goes wrong they rewind in text and try again — which is both a quality problem and a token-cost problem.

This page separates three things the retellings merge: what the lab **measured**, what it **claimed**, and what the coverage **extrapolated**. Only the third one is wrong, and it is the one being repeated.

## What is Bad Theory Labs, and what has it published?

A lab that publishes artifacts you can run. Two are relevant here, and they are different kinds of thing — which is the first place the coverage goes wrong.

**Interference Search** is an inference-time search method, not a model. Its code is public under **Apache-2.0** at [github.com/Badtheorylabs/interference-search](https://github.com/Badtheorylabs/interference-search), alongside the paper, the trained judge, the raw results, and a research log that includes failed experiments.

**BTL-4** is a 35B-parameter model published on Hugging Face under **Apache-2.0**, for tool use, software engineering and execution-grounded coding, with a native context of **262,144 tokens**. Its [model card](https://huggingface.co/badtheorylabs/BTL-4) states it is fine-tuned from Ornith-1.0-35B on an execution-gated reasoning corpus.

That last detail matters and is almost always dropped: **BTL-4 is a fine-tune, not a new architecture.** The architecture story is Interference Search, a separate artifact.

The companion runtime for the earlier [BTL-3](https://github.com/Badtheorylabs/BTL-3) is public under the **MIT licence**, while its model artifacts are Apache-2.0 — two licences covering two different things, which is the kind of distinction worth checking yourself rather than inheriting from a summary.

## Is there a paper behind the non-linear reasoning claim?

Yes — and this is the correction we had to make to our own draft. There is a paper, at [badtheorylabs.com/papers/interference-search](https://www.badtheorylabs.com/papers/interference-search) and as `paper/PAPER.pdf` in the repository, with the LaTeX source beside it.

It is **not on arXiv and not peer-reviewed**. Those are real limitations and worth stating. But "no paper, therefore unverifiable" — the framing we started with, and the framing much of the coverage implies — is false.

In one respect the release is **more** reproducible than a typical preprint. The repository carries `paper/CLAIMS.md`, a table mapping every number in the paper to the file that holds the result and the command that regenerates it.

It also ships the trained judge weights, the raw result JSON, and `docs/RESEARCH_LOG.md` including experiments that did not work. Publishing your failures is not what an unfalsifiable claim looks like.

The lab is also careful where it would be easy not to be. The repository states that the name comes from quantum search, where wrong paths cancel — and then says plainly that the method is classical and claims no quantum speedup.

## What did Interference Search actually measure?

Search over merged states instead of a single transcript. Every live branch expands at once, the environment executes the moves, branches reaching the same state merge, a small trained judge drops states that can no longer reach the goal, and survivors advance together.

The headline results, from the claims log:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Measurement</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">Interference Search</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">One line of thought</th>
    </tr>
  </thead>
  <tbody>
    <tr style="background:#f0f9ff; border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Hard 4-number Countdown solved, same judge</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>30/30</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">21/30</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">Sequential steps to solve all 30</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>3.0</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">23.7</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">State compression at six numbers</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;"><strong>13,229 states</strong></td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">831,176 paths</td>
    </tr>
  </tbody>
</table>

The three-step figure is where the widely-quoted "three steps" comes from, and it is a measurement with a command attached. The speed multipliers circulating alongside it are derived from the same runs rather than invented.

Two supporting findings explain *why* it works, and they are the most interesting part. In the first experiment, **45.1%** of one model's failed attempts repeated an expression it had already ruled out itself.

On another problem the model wrote the correct answer at token 1,313 and never committed to it. Merging duplicate states and pruning dead ones attacks exactly those two failures.

## Does a Countdown result predict enterprise agent reasoning?

That is the unsupported step, and it belongs to the retelling rather than to the lab. Countdown gives you a few numbers and a target and asks for an expression using each once.

It was chosen for good reasons the repository states openly: the arithmetic is trivial, the difficulty is purely in choosing which moves to make, and **it has an exact solver** — so every state can be labelled alive or dead without an LLM grading it.

That oracle is what makes the experiment clean. It is also what makes it unlike your workload.

An agent handling a refund policy, a prior-authorization appeal or a student's code has no exact solver. There is no function that labels a conversational state dead, which means the trained judge at the centre of the method has nothing to learn from in the same way.

None of that makes the result uninteresting — a method that removes redundant reasoning is worth watching, and the failure modes it targets are real in production.

It does mean **"cracked non-linear AI reasoning, implications for enterprise agents are massive" is a claim nobody has tested**, including the lab, which did not make it.

## What do BTL-4's benchmark numbers say?

They are specific, self-reported, and more modest than the architecture story. From the model card:

<table style="width:100%; border-collapse:collapse; margin:1.5rem 0; font-size:0.95rem;">
  <thead>
    <tr style="background:#f5f5f0; border-bottom:2px solid #2175C5;">
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Benchmark</th>
      <th style="text-align:right; padding:0.75rem; color:#5f6368;">BTL-4</th>
      <th style="text-align:left; padding:0.75rem; color:#5f6368;">Note</th>
    </tr>
  </thead>
  <tbody>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">SWE-bench Verified</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">78.4%</td>
      <td style="padding:0.75rem;">Self-reported</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">BFCL v4 (AST)</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">73.5%</td>
      <td style="padding:0.75rem;">+4.3 points over the 69.2% base</td>
    </tr>
    <tr style="border-bottom:1px solid #e5e7eb;">
      <td style="padding:0.75rem;">LiveCodeBench v6</td>
      <td style="text-align:right; padding:0.75rem; font-variant-numeric:tabular-nums;">66.1%</td>
      <td style="padding:0.75rem;">Easy 99.1% · Medium 86.7% · Hard 60.5%</td>
    </tr>
  </tbody>
</table>

One line on the card is more instructive than the rest. LiveCodeBench moved from **60.9% to 66.1% purely by raising the output budget from 16K to 32K tokens** — the model did not change, the allowance did.

That is useful honesty, and it complicates any efficiency story told around the same model. A score that improves when you let it write more is buying quality with tokens.

The spread from 99.1% on easy problems to 60.5% on hard ones is the other number to sit with. It is the shape of most coding results, and it is why a benchmark average says little about the tasks your engineers actually bring.

## How should an enterprise evaluate a claim like this?

Four checks, in order. They would have caught our own first draft.

**Look for the artifacts before judging the claim.** We assumed marketing and found a claims-to-evidence log. Absence of an arXiv link is not absence of evidence, and a self-published paper with runnable code is stronger than a preprint with neither.

**Read the benchmark, not the multiplier.** A figure is only as general as the task it was measured on. "30/30 versus 21/30 on hard Countdown with the same judge" is a precise statement; "4,200× faster" detached from Countdown is not a statement about your workload.

**Separate the artifacts.** A fine-tuned model, an inference-time search method and a runtime are three different things with three different claims. Coverage collapses them, and then the strongest claim gets attached to whichever one you were considering buying.

**Run your own set.** The only benchmark that decides anything is made of your tasks and policies. Nubank screened open-weight configurations across **more than 16,000 simulated conversations** before choosing one, which is what a defensible model decision looks like: [Nubank Screened 16,000 Simulated Chats Before Going Live](/blog/nubank-simulated-conversations-agent-testing-before-production).

## What would it cost you to test a new model next week?

That question decides whether any of this matters to you, and the answer depends entirely on what you bought.

If your AI platform is a per-seat subscription to a vendor's own models, the answer is close to "you cannot" — the model is the product, so an open-weight challenger is something you read about. You find out whether the method helps when your vendor adopts it, or does not.

If you own the stack, the answer is a few hours. Apache-2.0 weights can be pulled, served and pointed at your existing evaluation set, and if the numbers do not hold on your work you have spent an afternoon.

That is the practical content of being model-agnostic, and it is why we treat it as an architectural property rather than a feature. The same argument applies to the frontier labs' releases: [Frontier Labs Set the Pace](/blog/frontier-labs-pace-the-frontier-model-agnostic-infrastructure).

On ibl.ai you own all the code and the data, the platform is model-agnostic across any LLM, and there is no per-seat pricing — so swapping in a candidate is configuration rather than procurement.

The agent sandbox contains the test: code runs in a Linux virtual machine with **no network at all** by default, opened only to hosts you allowlist, so an unvetted model driving unvetted code stays bounded.

That mechanism is described in [Letting a K-12 AI Agent Run Code Without Letting Data Out](/blog/agent-sandboxes-k12-districts-code-execution).

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

The asymmetry is the point. An organization that can test an extraordinary claim in an afternoon never has to predict which claims are true.

## Want to screen a new model against your own workload?

We can stand up an evaluation set from your own tasks and policies and run any candidate model against it, on infrastructure you own. [Book a 30-minute demo](https://cal.com/iblai/30min) or [talk to the ibl.ai team](/contact) — ibl.ai is family-owned and operated from New York, NY.

## Why does owning the AI stack matter?

**ibl.ai is the agentic AI platform where you own all the code and the data. You self-host the entire stack inside your own perimeter, run it model-agnostic across any LLM and switch anytime, and pay by usage with no per-seat pricing — so you can deploy anywhere: your cloud, on-premise, GovCloud, or fully air-gapped.**

- **You own all the code and the data.** Full source code under a perpetual license, running on your infrastructure. Not API access to someone else's platform — the stack itself is yours.
- **Model-agnostic.** Run any LLM — Claude, GPT, Gemini, Llama, Command, or your own fine-tune — and switch providers without rewriting the platform.
- **No per-seat pricing.** Usage-based billing against a budget cap you set. Cost tracks what your organization actually uses, not how many people you employ.
- **Deploy anywhere.** Your cloud, your VPC, on-premise, GovCloud, or a fully air-gapped network with no outbound connectivity.

1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.

ibl.ai is family-owned and operated from New York, NY — a U.S.-headquartered, domestically-owned long-term partner, not a vendor that sells licenses and moves on.
