HOBO docs
What HOBO is
HOBO is a decision model. Give it some text, a question and a short list of options. It returns one option and how sure it is of each. It never writes text.
The idea
Most AI models answer every time, in the same confident voice, whether they're right or not. HOBO is built the other way: its confidence is meant to be used. You keep the answers it's sure of, it asks for a missing detail when one is missing, and everything else comes back to you.
On 1,933 real tool-call requests it never trained on, HOBO is right 88.8% of the time. With a bar fitted on 300 labelled requests it decides about 88.8% of them on its own, keeps those at 90% right or better, and hands the rest back. See Fit your bar.
What it decides
HOBO is trained for the decisions an AI agent makes all day. Each has its own page under Decisions.
| Decision | The question | The answers |
|---|---|---|
| Which tool | Which function fits this request? | a function, ask first, or none fits |
| A detail given | Does the request give this detail? | yes or no |
| Refusal | Did the assistant answer or refuse? | answered or refused |
| Claim | Does the passage back the claim? | supported, refuted, or not stated |
| Reading | Which part of the passage answers the question? | a span, or none |
| Intent | What does the message want? | one of your intents |
Other questions with a short list of answers work too, once you fit a bar on your own checked examples.
What it isn't
- Not a chatbot. It never writes. Every answer is one of the options you gave it.
- Not always right. It's built to know when it might be wrong, and to say so.
- Not a safety filter on its own. It judges refusals and requests well, but the policy is yours.
Our goal
An honest, unbiased decision machine: sure only as often as it's right, quick to hand back what it can't settle, and judging the evidence in front of it, not who it came from. Every version is tested against that goal before it ships, and the results are published as they come, in Where HOBO struggles too.
Where it comes from
HOBO starts from a publicly released base model, fine-tuned on 35,658 rows, every one listed with its source and license. The numbers here are from the current preview checkpoint, which passed every bar registered before its training run. More training may follow, each run registered before it starts. See How we test and Data and licensing. An earlier, smaller model, HOBO Mini, is archived and no longer developed.
Next
- Fit your bar: the one step that makes the confidence yours.
- Decisions: each kind of question, and how well HOBO does it.
- Patterns: hand back, ask first, and small decisions combined in code.
- Where HOBO struggles: the measured weak spots.