Archived, 2026-10-02. HOBO Mini is no longer developed. Its weights and this page stay up as a record. HOBO replaces it.
A small model that picks one answer from a short list, and says how sure she is.
HOBO gives you the answer and a number for how sure she is. You set a bar for what to keep; whatever falls under it goes to a person instead of a guess, the hand off. Verdict in one line: every row she learned from is listed, and she runs on your own machine.
of the time she can tell a refusal from an answer, on a test whose responses she never trained on.
right on 1,933 real tool calls she never trained on: which function fits, or none.
per answer on a laptop CPU, faster on any GPU. Runs on your machine; nothing leaves it.
She ships with a training manifest: one line for every row she could have trained on, with its source, its license, the model that wrote it where one did, and the model that checked it.
| Where the rows came from | Rows |
|---|---|
| Built by our own programs: puzzles, records, invented encyclopedia articles, where the right answer is computed by the code that asked the question | 59,800 |
| Written by open models whose terms allow training on their outputs (Qwen3, Mistral-Small, Granite under Apache 2.0; Phi-4 under MIT); 16,978 answered again, blind, by a second model. 4,774 are an older build of the reply set a training machine still held, marked in the manifest | 18,162 |
| From published datasets that ask only for credit: CLINC150, BANKING77, Mind2Web, TruthfulQA | 16,527 |
| Left out because they shared text with a test set; the manifest says which, and why. One more, found after training, is marked | 1,471 |
How we got here. We set aside a base model because its training included share-alike data, and another encoder checkpoint over books that a pending lawsuit alleges include Books3. When our own audit found passages from a test set inside one of our training sets, we corrected our records and rebuilt her from scratch.
What we can't tell you: her base encoder, ModernBERT-large (Apache 2.0), was pretrained on data its authors describe but do not itemise. That is true of nearly every model built on a public base, and it is the one layer we did not train. The audit tools are public: dinostomp checks the data, moons-dont-talk checks the training run.
The same items and the same question for every model: HOBO Mini, this page's model; HOBO, the full-size model; Laya, the open model whose head design HOBO's follows, as it ships; and HOBO's base, the model it starts from, before any HOBO training. Best in each row in bold.
| What we asked | HOBO Mini | HOBO | Laya | HOBO's base |
|---|---|---|---|---|
| Is this reply a refusal?XSTest frozen test, 255 · HOBO and HOBO Mini trained on XSTest prompts, never these replies | 96.5% | 97.3%94.4–98.7 | 74.9% | 37.6% |
| Which function fits this request, or none?BFCL v4 live, 1,933 real requests · never trained on | 78.1% | 88.8%87.3–90.2 | 38.3% | 21.9% |
| Which span of the passage answers the question?SQuAD 2.0 passages, 1,000 · never trained on SQuAD; its question format is in training since this preview | 89.5% | 93.6%91.9–95.0 | 49.3% | 68.5% |
| Does the evidence support or refute the claim?FEVER, 1,000 · never trained on FEVER | 79.1% | 81.4%78.9–83.7 | 85.5% | 78.2% |
| Which intent does the message express?CLINC150, 1,000 · HOBO and HOBO Mini trained on its training split | 96.2% | 95.2%93.7–96.4 | 92.3% | 42.5% |
| Which banking intent is this?BANKING77, 1,000 · HOBO and HOBO Mini trained on its training split | 89.3% | 84.4%82.0–86.5 | 69.7% | 26.2% |
HOBO is in preview: three training runs whose answers are averaged, with a 95% interval under each score. It is built for the decisions an agent makes (refusals, tool calls, claims); both HOBOs practised on the two intent sets, and Mini still leads there. Its clearest gain is routing tool calls. HOBO's base, before any HOBO training, sits below the 45.5% you'd get by always answering "none" there, so that gain is HOBO's; on FEVER, its base already reads 78.2%. Laya was read as it ships, with the same questions HOBO gets; its own prompts may suit it better. "Never trained on" means HOBO saw zero rows from that set.
from hobo_preview import load
model = load(".")
d = model.refusal("Sorry, I can't help with that.")
d.said, d.p # ("refusal", 1.0)
keep = d.p >= bar # bar: fitted on your own labelled rowsThe full-size HOBO is aimed at the decisions an AI agent makes all day: whether a reply refuses or answers, which tool to call, or none, or ask for more detail first, and whether a passage backs a claim. Same rules as this one: every training set is audited before she sees it, and she'll ship with every row she learned from listed.
We're also teaching her Korean. We'll say she reads it once a test we didn't write says so. HOBO is already in the table above. The model on this page stays as HOBO Mini: five times smaller, for machines where every second counts.
The research preview is coming to Hugging Face: the weights, a small loader with one method per task, the model card, and the training manifest. Evaluation and research use only. Until then, ask@collapseindex.org.
HOBO is a research preview, not the release. The numbers on this page are from the checkpoint of record, read on held-out sets; where a set was never trained on, the page says so. Nothing leaves the machine; gum not included.
Alex Kwon · ask@collapseindex.org· more models· collapseindex.org