newcombsA research lab for decisions a machine has to commit to.
Software often has to commit to a small decision: which queue a ticket goes to, whether it is urgent, whether a person should see it today. Chat models answer by writing text, and the program then has to find the label in that text. In this study the model returns a probability for each answer listed in advance.
The lab is named after Newcomb's problem, a puzzle about a predictor that has already guessed your choice. Typed decisions are the current study. You send a situation as JSON plus a few questions. Each question is yes/no, one of several options, or a score on a scale.
The model writes no prose. Every option is scored against the situation and the scores are turned into probabilities, so there is nothing to parse afterwards.
Onebox 1 is out. It is our first open model, 9 billion parameters, on Hugging Face at newcombs/onebox-1-9b with weights and inference code under the Apache 2.0 license.
Two boxes are on the table. You can see that one holds $1,000. The other is closed. It holds either a million dollars or nothing.
A predictor has already guessed what you will do. If it guessed you would take only the closed box, it put the million in. If it guessed you would take both, it left the closed box empty. The guess is made. Taking a box does not change what is already inside.
Take both, says one argument. Whatever is in the closed box, the clear box adds $1,000. If the million is there, you leave with a million plus $1,000. If the closed box is empty, you leave with $1,000 instead of nothing. This is called two-boxing. The theory behind it is causal decision theory: choose the act that causes the better result, given how the boxes already are.
Take only the closed box, says the other argument. In the story the predictor is very accurate. People who take only the closed box tend to find the million. People who take both tend to find it empty. The $1,000 you can see is how you miss the million. This is called one-boxing. The theory behind it is evidential decision theory: your choice is evidence of what the predictor already did, so you pick the choice that an accurate predictor would have rewarded.
The two arguments disagree, and each one looks finished. Robert Nozick wrote the problem up in 1969, from a puzzle by the physicist William Newcomb. Decision theory has not settled it.
The lab takes its name from that predictor, who has to commit before you choose. Our models are called Onebox, after the choice to take one box. They also decide in one pass.
Onebox 1 (9B) is built on Qwen3.5-9B. We trained 44 million parameters on top of it, in about two hours on one rented GPU. It runs on a Mac, on an NVIDIA GPU or on the CPU.
pip install torch "transformers>=5.10" safetensors huggingface_hub
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("newcombs/onebox-1-9b")
sys.path.insert(0, path)
from onebox import load
model = load(path)
answer = model.decide(state, questions)
The model card on Hugging Face has the full numbers, the limits and how we picked this model from seven trained variants.
A request looks like this:
{
"state": {
"ticket": "Server down since 3am, customers cannot log in"
},
"questions": {
"team": {
"type": "choice",
"instructions": "Which team handles this?",
"criteria": {
"billing": "payments, invoices",
"technical": "outages, bugs, login",
"sales": "new contracts"
}
},
"urgent": {
"type": "noul",
"instructions": "Does this need action today?"
}
}
}
And what Onebox 1 answers, rounded:
{
"answers": {
"team": {
"type": "choice",
"choice": "technical",
"probabilities": {"billing": 0.035, "technical": 0.931, "sales": 0.034}
},
"urgent": {"type": "noul", "noul": 0.977}
}
}
We use the same request format as TypeSafe's System One API, so requests written for it work with Onebox too.
We test on the public benchmark LocalLLaMA/typed-decisions, official test split: 400 cases with 2000 decisions from four workflows (agent traces, customer service, invoices, security incidents). A decision counts as correct if our most likely option matches the label. To make sure our setup is fair, we ran Laya ourselves on the same split and got its published numbers back to the third decimal.
| model | trained on these four workflows? | accuracy |
|---|---|---|
| always the most common answer | n/a | 46.1 % |
| Laya 421M | no | 36.2 % |
| our recipe without the training split | no | 59.4 % |
| openjev 4B v5 | no | 63.8 % |
| Clef-Flash 9B | not as far as we know | 70.3 % |
| Julia-1 | unknown | 72.6 % |
| TypeSafe Jev (closed, published number) | unknown | 72.7 % |
| Laya 421M, fine-tuned | yes | 76.6 % |
| Onebox 1 (9B) | yes (5370 decisions) | 79.3 % |
The 79.3 % needs a caveat. Onebox saw the training split of the same four workflows, so this is the number for work it has seen before. The same recipe without that split gets 59.4 %, which is closer to what to expect on a new workflow. Whether training on many more kinds of decisions closes that gap is the question we work on.
A lure is a word in the input that points toward the wrong answer. A customer writes "I can't open the billing page." The page does not load, so this is a job for the technical team, but the word "billing" pulls toward billing.
Onebox 1 falls for this one with 99.7 %. So did Laya (99.9 %) and Julia-1 (98.9 %). Clef-Flash was the only model we tried that picked technical, with 70 %.
To see how often this happens, we built two sets of 200 pairs. In each pair one case has a lure and the other uses the same word where it really points to the right answer. A pair counts when both cases are right. We never train on these sets.
| model | set A | set B |
|---|---|---|
| Onebox 1 (9B) | 85.0 % | 60.5 % |
| Clef-Flash 9B | 90.5 % | 69.5 % |
Clef-Flash is built on the same base model and does better, so this is not a limit of the base model. Our first attempt to fix it is better training data, the main goal for Onebox 1.1.
Onebox 1 is public: weights and inference code are on Hugging Face. We keep a lab notebook where every number on this page is written down along with how we measured it, failed runs included.
Next comes Onebox 1.1, trained to stop falling for lures, and a smaller version that runs well on a laptop without a big GPU.
newcombs, Berlin. Last updated 5 October 2026.