HOBO docs

Data and licensing

A source behind every training row. Here's where they come from and what we leave out.

Every row, listed

HOBO trained on 35,658 rows. The release comes with a manifest listing each one with its source, its license, and, where a model wrote it, which model wrote it and which checked it. Some rows are variants our code made of other training rows (the same text with a pushback note, a source line or a quote added); each is listed with the row it came from and keeps that row's source and license.

Where the rows come fromRows
Human-written public datasets8,377
Built by code (articles about invented places and people, effect-size reports)8,000
Written by open models whose licenses allow training on their outputs19,281
Total35,658

Sources and licenses

  • Base model: HOBO starts from a publicly released base model. Its makers describe its pretraining data but don't itemise it, like nearly every public base model, so that part we can't list.
  • Datasets: human-written datasets that ask only for credit or are MIT-licensed: CLINC150, BANKING77, XSTest's prompts, MASSIVE, SciFact, Mind2Web, MultiWOZ.
  • Model-written text: only from open models released under Apache 2.0 or MIT. A second open model checks each label blind.

What we leave out

  • Non-commercial data, share-alike data, and outputs from closed models.
  • Test sets. SQuAD, FEVER, BFCL and Do-Not-Answer are only ever used to measure.
  • Rows too close to a test. 360 training rows that shared text with a test set were taken out before training.

Audited first

Every dataset goes through dinostomp, our open dataset audit, before it's trained on or measured against: duplicates, leaks between splits, label skew, answers given away by wording.