HOBO docs
Data and licensing
A source behind every training row. Here's where they come from and what we leave out.
Every row, listed
HOBO trained on 35,658 rows. The release comes with a manifest listing each one with its source, its license, and, where a model wrote it, which model wrote it and which checked it. Some rows are variants our code made of other training rows (the same text with a pushback note, a source line or a quote added); each is listed with the row it came from and keeps that row's source and license.
| Where the rows come from | Rows |
|---|---|
| Human-written public datasets | 8,377 |
| Built by code (articles about invented places and people, effect-size reports) | 8,000 |
| Written by open models whose licenses allow training on their outputs | 19,281 |
| Total | 35,658 |
Sources and licenses
- Base model: HOBO starts from a publicly released base model. Its makers describe its pretraining data but don't itemise it, like nearly every public base model, so that part we can't list.
- Datasets: human-written datasets that ask only for credit or are MIT-licensed: CLINC150, BANKING77, XSTest's prompts, MASSIVE, SciFact, Mind2Web, MultiWOZ.
- Model-written text: only from open models released under Apache 2.0 or MIT. A second open model checks each label blind.
What we leave out
- Non-commercial data, share-alike data, and outputs from closed models.
- Test sets. SQuAD, FEVER, BFCL and Do-Not-Answer are only ever used to measure.
- Rows too close to a test. 360 training rows that shared text with a test set were taken out before training.
Audited first
Every dataset goes through dinostomp, our open dataset audit, before it's trained on or measured against: duplicates, leaks between splits, label skew, answers given away by wording.