On September 18, Bespoke Labs released Nimble, an open weight model that makes typed decisions: you hand it some state and a set of questions with fixed answer options, and it hands back a choice and a probability. No generated text, nothing to parse. It is the same job that Jev, the closed model TypeSafe launched three days earlier, was built to do.
The claim going round since is that Nimble does that job just as well as Jev. It does not, quite. It is close, closer than anyone expected from a single day of training, and it is the first model of this kind you can actually download and run on your own hardware. Those are two different things, and the second one is the interesting one.
We went through the benchmark tables in the repository rather than the screenshots, because the two tell somewhat different stories.
This post covers:
- What a decision model does that a chat model does not
- What Bespoke actually trained, and how quickly
- The 90 percent figure everyone is quoting, and which test it comes from
- The harder set of public benchmarks, where the gap looks different
- Calibration, which is where Jev still leads clearly
- What it takes to run Nimble yourself, and what that costs
- Eight places a decision model earns its keep
- The limits worth knowing before you build on it
1. What a decision model does that a chat model does not
Most teams get structured output from a language model by asking for JSON and validating what comes back. The model still generates tokens, you still parse them, and it can still produce a field you never defined. Retries and validators exist because that path fails often enough to need them.
A decision model skips generation. You give it the state (a passage, a ticket, a conversation) and a schema of questions. It scores only the answers you allowed and returns one of them with a probability attached. Ask it to pick among three tools and it returns one of those three, because a fourth was never a candidate. Both Jev and Nimble handle three question shapes:
| Question shape | What you get back | Typical use |
|---|---|---|
| Choice | One option from a fixed list, Nimble allows up to 26 per field | Intent routing, tool selection, entailment |
| Boolean (Jev calls it a Noul) | True or false with a probability | Policy gates, answerability, relevance |
| Score | A level on an ordered rubric, plus a probability weighted average | Severity, quality rating, confidence tiers |
The practical difference is where these things sit. A chat model sits at the end of a pipeline producing something a person reads. A decision model sits in the middle of one, in the hot path, called many times per task, deciding what happens next.
2. What Bespoke actually trained
The recipe is unusually plain, which is part of why it landed so fast.
- Base model: Qwen3.5-9B
- Method: a LoRA adapter at rank 16, one epoch, learning rate 5e-5, batch size 8. The adapter is about 165 MiB.
- Training data: 2,676 synthetic examples across ten subject categories
- License: Apache 2.0, with the weights, the data, and the recipe all published
The part worth attention is what Bespoke calls contrastive data curation. Rather than labelling probabilities directly, they generate near duplicate examples where a small change to a fact flips the correct answer, and train the model to separate them. Calibration falls out of the data design instead of being fitted afterward. Bespoke also state that they did not distil Jev, and used it as an evaluator rather than a teacher, which matters if you care about the license position of what you deploy.
Training time was roughly a day. Bespoke are upfront that the project is young and has rough edges, and nothing below contradicts that.
3. The 90 percent figure, and which test it comes from
This is the table in the tweet that started the discussion:
| Model | Agreement |
|---|---|
| Jev 1.13.0 | 93.21% |
| Bespoke-Nimble-9B | 90.12% |
| Qwen3.8-27B, untuned | 84.88% |
| Qwen3.5-9B, the base model | 66.36% |
Two readings of that. The impressive one is the bottom row: the same base model went from 66% to 90% on a day of training against 2,676 examples. That is a real result about how tractable this task is.
The cautious one is the test itself. It is Nimble's own 324 example holdout, drawn from six source families, with synthetic labels that were checked by a model and by no person. Bespoke say so plainly in their documentation. A holdout built by the same pipeline as the training data tells you the model learned the pipeline. It does not tell you much about what happens on somebody else's problem, and it is not a standardised benchmark that anyone else reports against.
4. The harder set
To their credit, Bespoke also ran both models over thirteen public subsets with human written labels, 3,880 records in total, covering intent routing, entailment, paraphrase detection, moderation, guardrails, medical question answering, and rubric rating. None of those categories were in Nimble's training data. Here is what came out:
| Group | Subsets | Nimble macro | Nimble micro | Jev macro | Jev micro |
|---|---|---|---|---|---|
| All | 13 | 74.8% | 75.9% | 76.0% | 77.3% |
| Choice | 5 | 81.6% | 81.1% | 82.9% | 82.8% |
| Boolean | 5 | 80.2% | 80.1% | 84.6% | 84.6% |
| Score | 3 | 54.6% | 51.2% | 50.1% | 45.2% |
A gap of 1.2 macro points overall, which is narrower than the holdout suggested. The shape underneath is more useful than the average. Jev leads clearly on the boolean tasks, by about four points, and slightly on choice tasks. Nimble leads on the rubric scoring tasks, where both models are weak in absolute terms.
Five of the thirteen per subset gaps are statistically separable. Four favour Jev (Civil Comments, PAWS, German MASSIVE, and VitaminC) and one favours Nimble (SummEval relevance). The other eight do not separate the two models on this evidence, which is a polite way of saying the sample sizes are a few hundred records each and you should not read much into a two point difference.
One caveat applies to the whole suite and is easy to miss: these numbers measure agreement with human annotators, not correctness. Annotators disagree with each other, and guidelines differ between projects.
5. Calibration is where the gap is real
If you only take one number from this post, do not take the accuracy. Take the calibration.
A decision model's probability is the thing you build control flow on. You escalate to a person below 0.8, you auto approve above 0.95, you retry in between. That only works if 0.9 actually means nine times in ten.
Jev has the lower expected calibration error on 11 of the 13 subsets and the lower Brier score on 10 of 13. Neither model has had a temperature fitted, so this describes them as shipped. On the two subsets that carry a full distribution of human votes the two are close, so the gap is not uniform, but the direction is consistent.
There is a second pattern worth knowing if you work in more than one language. Bespoke ran the same 350 utterances through MASSIVE in English and in German. Nimble lost 3.5 points switching language. Jev lost 0.5.
| Jev 1.13.0 | Bespoke Nimble 9B | |
|---|---|---|
| Overall public accuracy | 76.0% macro | 74.8% macro |
| Boolean questions | Leads by about 4 points | Behind |
| Rubric scoring | Behind | Leads by about 4 points |
| Calibration (expected calibration error) | Better on 11 of 13 subsets | Better on 2 |
| Non English | Holds up, 0.5 point drop | 3.5 point drop |
| Weights | None, closed API | Apache 2.0, downloadable |
| Schemas | Nested structures supported | Flat only |
| Where it runs | TypeSafe, OpenRouter, Vercel, Cloudflare | Anywhere you have a GPU |

6. What it takes to run it
Nimble is a 9B model. At BF16 the weights take about 18 GB, so it fits on a single 48 GB card with a lot of room left over, which turns out to matter (more on that at the end). Bespoke published latency from two machines:
| Hardware | Median | Mean | p95 |
|---|---|---|---|
| NVIDIA H100 | 106 ms | 110 ms | 120 ms |
| Apple M5 Pro, 64 GB, MLX | 444 ms | 546 ms | 981 ms |
For comparison, in Bespoke's own evaluation runs Jev answered through its hosted API at a median of about 660 ms per request including network transport. That is not a serving benchmark for either model, and Bespoke say as much, but it points at the argument that actually holds up for running this locally.
Because the cost argument does not. Jev is cheap: roughly $0.042 per million input tokens with output free. A decision heavy agent making twenty calls per task at 600 tokens each would run about $500 a month at a million tasks. That is not the kind of bill that justifies a GPU on its own.
What justifies it is everything else:
- Latency in the hot path. Decisions are called many times per task. A round trip to somebody else's API, repeated twenty times, is the difference between an agent that feels immediate and one that does not.
- The data never leaves. More on this below, because it is the case we find most convincing.
- No waitlist and no version drift. Jev is early access. Nimble is a file you pin, run, and keep running when the vendor deprecates something.
- Nobody hosts it for you. As of writing, no inference provider serves Bespoke Nimble. Running it means running it yourself, on your hardware or somebody's dedicated hardware.
Two configuration details will bite if you skip the documentation. Nimble was trained with a 2,048 token budget. The serving limit was recently raised to 8,192 tokens, configurable through NIMBLE_MAX_PROMPT_TOKENS, but the authors are explicit that quality above the training budget has not been established. Longer prompts are accepted, not validated. And schemas must be flat, with at most 26 options per field.
7. Eight places a decision model earns its keep
We have been asked about most of these by customers running agents on private hardware. They share a shape: high call volume, a small fixed answer space, and a decision that something downstream depends on.
- Tool and handler routing. Pick one of N tools for this request. The single most common use, and the one where never inventing an option matters most, because an invented tool name is a crash rather than a bad answer.
- Loop termination. Asking "is this task finished?" after every agent step. Cheap to ask, expensive to get wrong in both directions, and a chat model asked the same question tends to be agreeable.
- Retrieval gating. Before a passage goes into a prompt, ask whether it actually answers the question. This trims context, cuts cost on the expensive model downstream, and reduces the confident wrong answers that come from padding a prompt with near misses.
- Citation checking. For each claim and source pair, a three way choice: supports, contradicts, or does not address. This is the verification step most retrieval systems skip, and it is a near perfect fit for a typed choice.
- Guardrails and policy gates. One boolean per policy, evaluated before a prompt reaches a model or an action executes. Nimble scored 81.2% on the Aegis content safety set, slightly ahead of Jev, though the difference is not separable.
- Ticket triage and severity scoring. Category as a choice, severity as a rubric score. Rubric scoring is the one family where Nimble leads, which makes this a reasonable place to start if you want to trial it.
- Extraction validation. After a pipeline pulls fields from a document, ask whether each extracted value is actually supported by the source. Errors get caught before they reach a database instead of after.
- Sensitivity classification at the boundary. This is the one we would build first, and it is worth spelling out.
Plenty of teams run a hybrid setup: a hosted frontier model for general work, a private stack for anything confidential. Something has to decide which is which, document by document. If that classifier is itself an API call, then every document gets sent to a vendor in order to determine whether it was allowed to be sent to a vendor. The gate has to sit inside the boundary it is guarding, which means it has to be a model you can run yourself. Until this month there was no good small model built for that job. Now there is one.
8. Limits worth knowing before you build on it
Bespoke document most of these, and we would rather repeat them than have somebody discover them in production:
- Text only. No images, no audio.
- Flat schemas only. Nested question structures, which Jev handles, are not supported.
- The 2,048 token training budget described above. Long documents need chunking, and chunking changes the question you are asking.
- 26 options per field, which is generous for routing and tight for classification over a large taxonomy.
- Probabilities are not correctness. A confident wrong answer is still available to you, and the calibration numbers say it is somewhat more available here than with Jev.
- Weaker outside English, on the one paired test that exists.
- It is days old. The benchmarks are the authors' own, and no independent party has reproduced them yet.
For agent routing, policy gates, completion checks, and severity scoring, none of that is disqualifying. For decisions where a miscalibrated probability carries regulatory or financial weight, Jev is still the better engineered product, and saying otherwise would not survive contact with the tables above.
9. The other repository people are mixing up
A second project, openJev-verdict-2.0, has been circulating in the same threads and its numbers are being quoted next to Nimble's. They are not comparable. It is a 149.6M parameter ModernBERT encoder, non autoregressive, running in a browser through WebGPU in about 600 MB and answering in 20 to 25 ms. It reports 77.1% against Jev's 72.7%, but on a different 2,000 example set of typed decisions, not the one Nimble was measured on.
Different architecture, different test, different scale. It is a genuinely interesting option for edge and in browser work where 600 MB and no server is the whole point. It is not evidence about Nimble, and the 77% and the 90% should never appear in the same sentence.
Where this lands
Nimble is about three points behind Jev on its own test, one to two points behind on public sets, and clearly behind on calibration, after one day of LoRA training on a 9B chat model. Read as a claim of parity, that is overstated. Read as a measurement of how fast a capability moves from closed to open once somebody demonstrates it is possible, it is the more striking of the two stories, and it is the one we would pay attention to.
If you are building agents on infrastructure you control, it is the decision model you can run today.
Where Bit Refinery fits
We host GPUs, we do not build models. The reason this release is relevant to us is the sizing. A 9B decision model at BF16 needs about 18 GB, which leaves roughly 30 GB free on a single 48 GB RTX 6000 Ada. That is enough to run a quantised generative model on the same card, so the decision layer and the generation layer share one GPU and the decisions never cross a network at all. We wrote about how much work fits on one 48 GB card in the serving layer versus the agent harness, and about the same pattern with a larger open weight release in Kimi K3 on private GPU hosting.
If you want to try it on dedicated hardware, Private GPU Cloud gives you an NVIDIA GPU that serves you alone, with root access and no egress fees. BYOGPU is for cards you already own. Private LLM Hosting is for teams that would rather we ran the serving stack too. There is no model on our side of the rack, we do not log prompts, and we do not train on anything you run.
