Tools

Things you can run, open or try. Each one keeps its own defects on the record.

Models

  • HOBO Mini (archived)
    a small model that picks one answer from a short list, and says how sure it is.give it some text and a handful of options: is this reply a refusal, which tool fits this request, which passage answers the question, does the article support the claim. it returns one option and a probability for each, on your own machine. every row it trained on is listed with its source and license: data we generated, or data that asks only for credit. no SQuAD, no FEVER, and the one layer we did not train, the base encoder, is said plainly. 96.5% judging refusals, 78.1% routing live tool calls it never trained on. a research preview, archived on 2026-10-02 and no longer developed; HOBO replaces it.
    Alex Kwon · 2026 · 421M, runs local · the archived page

Tools

  • dinostomp
    a verification layer for AI evaluations.eval scores get published; the instrument that produced them almost never gets checked. dinostomp is that check, across the whole pipeline: the dataset, the scorer, the runs, whether the number survives its own sampling noise, and whether the claim has the evidence it needs. pointed at MMLU, it found a subtraction item keyed to two identical correct options. the error class was already known from MMLU-Redux, which hand-annotated 5,700 items; run against those annotations, dinostomp found two more double-keyed items the annotators had marked ok. prior art is cited in the ledger. it also audits agents: their tool traces are recorded by the harness rather than self-reported, and it can withhold an agent's retrieved evidence to ask whether the answer actually depended on it.
    Alex Kwon · 2026 · every check negative-tested against a planted defect · code
  • moons-dont-talk
    a preflight gate for ML training runs.moons don't talk; receipts do. before the GPUs get a job, it reads the data, the config, and the machine, runs a short canary of the real training command, proves the checkpoint by restoring it in a fresh process, projects runtime and cost against the limits you declared, and writes a receipt that binds the verdict to hashes of the data, config, code, and environment. change any of them and the verdict is invalid, field by field. any trainer that can honor five environment variables can be gated; the hugging face adapter has run on cpu, a t4, and an a10g. its first day of live tests produced six ledger entries against itself, including a disk check that passed on eight exabytes.
    Alex Kwon · 2026 · sibling of dinostomp: same method, different cargo · code

UX Prototypes

  • The Menu Was the Hard Part
    a working interaction prototype, not a mockup: real gesture, real drag physics, and no permission dialog it does not need.every chat app buries the camera behind a plus-menu and a file picker, then hands you a keyboard, which is fine at a desk and useless in a grocery aisle where one hand is holding the thing you are asking about. here the camera grows out of the ask bar itself on a swipe, a tap captures, and the suggested questions sit where a small model reading the frame would put them, so the answer is one more tap. the page contains both routes so the difference is something you feel: tap + for the conventional path and it counts your taps (five, plus typing) against the swipe (one swipe, two taps). typing never stops working, and both routes land in the same chat thread.
    Alex Kwon · 3 September 2026 · try it · code