The Short Answer
Bad Theory Labs published Interference Search with a paper, Apache-2.0 code, trained judge weights, raw results and a claims-to-evidence log, so its figures are measurements β measured on Countdown, an arithmetic puzzle with an exact solver. On ibl.ai you own all the code and the data and run the platform model-agnostic across any LLM, so testing whether a new method helps your own workload is an afternoon inside your own perimeter rather than a guess from someone else's benchmark.
The critique behind the work is real. Language models reason in a line, and when a step goes wrong they rewind in text and try again β which is both a quality problem and a token-cost problem.
This page separates three things the retellings merge: what the lab measured, what it claimed, and what the coverage extrapolated. Only the third one is wrong, and it is the one being repeated.
What is Bad Theory Labs, and what has it published?
A lab that publishes artifacts you can run. Two are relevant here, and they are different kinds of thing β which is the first place the coverage goes wrong.
Interference Search is an inference-time search method, not a model. Its code is public under Apache-2.0 at github.com/Badtheorylabs/interference-search, alongside the paper, the trained judge, the raw results, and a research log that includes failed experiments.
BTL-4 is a 35B-parameter model published on Hugging Face under Apache-2.0, for tool use, software engineering and execution-grounded coding, with a native context of 262,144 tokens. Its model card states it is fine-tuned from Ornith-1.0-35B on an execution-gated reasoning corpus.
That last detail matters and is almost always dropped: BTL-4 is a fine-tune, not a new architecture. The architecture story is Interference Search, a separate artifact.
The companion runtime for the earlier BTL-3 is public under the MIT licence, while its model artifacts are Apache-2.0 β two licences covering two different things, which is the kind of distinction worth checking yourself rather than inheriting from a summary.
Is there a paper behind the non-linear reasoning claim?
Yes β and this is the correction we had to make to our own draft. There is a paper, at badtheorylabs.com/papers/interference-search and as paper/PAPER.pdf in the repository, with the LaTeX source beside it.
It is not on arXiv and not peer-reviewed. Those are real limitations and worth stating. But "no paper, therefore unverifiable" β the framing we started with, and the framing much of the coverage implies β is false.
In one respect the release is more reproducible than a typical preprint. The repository carries paper/CLAIMS.md, a table mapping every number in the paper to the file that holds the result and the command that regenerates it.
It also ships the trained judge weights, the raw result JSON, and docs/RESEARCH_LOG.md including experiments that did not work. Publishing your failures is not what an unfalsifiable claim looks like.
The lab is also careful where it would be easy not to be. The repository states that the name comes from quantum search, where wrong paths cancel β and then says plainly that the method is classical and claims no quantum speedup.
What did Interference Search actually measure?
Search over merged states instead of a single transcript. Every live branch expands at once, the environment executes the moves, branches reaching the same state merge, a small trained judge drops states that can no longer reach the goal, and survivors advance together.
The headline results, from the claims log:
| Measurement | Interference Search | One line of thought |
|---|---|---|
| Hard 4-number Countdown solved, same judge | 30/30 | 21/30 |
| Sequential steps to solve all 30 | 3.0 | 23.7 |
| State compression at six numbers | 13,229 states | 831,176 paths |
The three-step figure is where the widely-quoted "three steps" comes from, and it is a measurement with a command attached. The speed multipliers circulating alongside it are derived from the same runs rather than invented.
Two supporting findings explain why it works, and they are the most interesting part. In the first experiment, 45.1% of one model's failed attempts repeated an expression it had already ruled out itself.
On another problem the model wrote the correct answer at token 1,313 and never committed to it. Merging duplicate states and pruning dead ones attacks exactly those two failures.
Does a Countdown result predict enterprise agent reasoning?
That is the unsupported step, and it belongs to the retelling rather than to the lab. Countdown gives you a few numbers and a target and asks for an expression using each once.
It was chosen for good reasons the repository states openly: the arithmetic is trivial, the difficulty is purely in choosing which moves to make, and it has an exact solver β so every state can be labelled alive or dead without an LLM grading it.
That oracle is what makes the experiment clean. It is also what makes it unlike your workload.
An agent handling a refund policy, a prior-authorization appeal or a student's code has no exact solver. There is no function that labels a conversational state dead, which means the trained judge at the centre of the method has nothing to learn from in the same way.
None of that makes the result uninteresting β a method that removes redundant reasoning is worth watching, and the failure modes it targets are real in production.
It does mean "cracked non-linear AI reasoning, implications for enterprise agents are massive" is a claim nobody has tested, including the lab, which did not make it.
What do BTL-4's benchmark numbers say?
They are specific, self-reported, and more modest than the architecture story. From the model card:
| Benchmark | BTL-4 | Note |
|---|---|---|
| SWE-bench Verified | 78.4% | Self-reported |
| BFCL v4 (AST) | 73.5% | +4.3 points over the 69.2% base |
| LiveCodeBench v6 | 66.1% | Easy 99.1% Β· Medium 86.7% Β· Hard 60.5% |
One line on the card is more instructive than the rest. LiveCodeBench moved from 60.9% to 66.1% purely by raising the output budget from 16K to 32K tokens β the model did not change, the allowance did.
That is useful honesty, and it complicates any efficiency story told around the same model. A score that improves when you let it write more is buying quality with tokens.
The spread from 99.1% on easy problems to 60.5% on hard ones is the other number to sit with. It is the shape of most coding results, and it is why a benchmark average says little about the tasks your engineers actually bring.
How should an enterprise evaluate a claim like this?
Four checks, in order. They would have caught our own first draft.
Look for the artifacts before judging the claim. We assumed marketing and found a claims-to-evidence log. Absence of an arXiv link is not absence of evidence, and a self-published paper with runnable code is stronger than a preprint with neither.
Read the benchmark, not the multiplier. A figure is only as general as the task it was measured on. "30/30 versus 21/30 on hard Countdown with the same judge" is a precise statement; "4,200Γ faster" detached from Countdown is not a statement about your workload.
Separate the artifacts. A fine-tuned model, an inference-time search method and a runtime are three different things with three different claims. Coverage collapses them, and then the strongest claim gets attached to whichever one you were considering buying.
Run your own set. The only benchmark that decides anything is made of your tasks and policies. Nubank screened open-weight configurations across more than 16,000 simulated conversations before choosing one, which is what a defensible model decision looks like: Nubank Screened 16,000 Simulated Chats Before Going Live.
What would it cost you to test a new model next week?
That question decides whether any of this matters to you, and the answer depends entirely on what you bought.
If your AI platform is a per-seat subscription to a vendor's own models, the answer is close to "you cannot" β the model is the product, so an open-weight challenger is something you read about. You find out whether the method helps when your vendor adopts it, or does not.
If you own the stack, the answer is a few hours. Apache-2.0 weights can be pulled, served and pointed at your existing evaluation set, and if the numbers do not hold on your work you have spent an afternoon.
That is the practical content of being model-agnostic, and it is why we treat it as an architectural property rather than a feature. The same argument applies to the frontier labs' releases: Frontier Labs Set the Pace.
On ibl.ai you own all the code and the data, the platform is model-agnostic across any LLM, and there is no per-seat pricing β so swapping in a candidate is configuration rather than procurement.
The agent sandbox contains the test: code runs in a Linux virtual machine with no network at all by default, opened only to hosts you allowlist, so an unvetted model driving unvetted code stays bounded.
That mechanism is described in Letting a K-12 AI Agent Run Code Without Letting Data Out.
1.6M+ users across 400+ organizations run the platform this way, including NVIDIA, MIT, and Syracuse University.
The asymmetry is the point. An organization that can test an extraordinary claim in an afternoon never has to predict which claims are true.
Want to screen a new model against your own workload?
We can stand up an evaluation set from your own tasks and policies and run any candidate model against it, on infrastructure you own. Book a 30-minute demo or talk to the ibl.ai team β ibl.ai is family-owned and operated from New York, NY.