← All work

Loom

Literature-based discovery, measured against elapsed time

Last devlog: July 28, 2026

DETERMINISTIC — THE SEARCH PROBABILISTIC — OFFLINE THE MODEL NEVER SEARCHES Corpus Windowliterature up to afixed past datenothing after it existsPath Searchmechanistic bridges,disease → substancethe expensive stageAdmissionwhich candidates thescorer may seebecame the failureSeam Scoredisjoint · distant ·multiply-bridgedranked top ten onlyBlind Graderdid this pair join afterthe corpus ended?block label strippedMatched Controldrawn from the engine's ownadmitted pool, matched onhow common each candidate isVerdictone answer per pair,every citation checked541 of 542 refusedA model asked to find connections invents them fluently — confident, unsupported claims on 63% of pairs. Here it cannot look anything up, and cannot tell ranked output from control. The boundary is what makes the answer mean something.

Loom is a literature-based discovery engine built on Don Swanson's 1980s premise: novel, true findings sit fully formed in published work, unnoticed because they span two fields that never cite each other. It searches the biomedical literature for mechanistic paths between a disease and substances never applied to it — deterministic search over a corpus windowed to a fixed past date, with the model confined to blind grading, never to searching.

The public result is a negative one, established honestly. The generative step was tested twice against bars frozen before any number existed, and failed both times. What makes the failure worth publishing: three separate causes were located to specific pipeline stages with numbers attached, and the matched control — drawn from the engine's own admitted pool — demonstrably could produce a positive, and did: a validated join the ranked surface missed. A negative result whose control can never score is unreadable; this one isn't.

The durable artifact is the verification apparatus: a grader that refused 541 of 542 fluent, confident, unsupported model claims; a pre-registration discipline that survived four review passes across three model families; and a closed arc whose every null has an address. The methodology is the part worth stealing.

Devlog posts about Loom

Believed, Not Verified

Five times in one project, a guarantee we were relying on was not in effect. A structured-output constraint silently ignored. A measurement that skipped a pipeline stage. A fix in the commit message but not the file. An exclusion in the specification but not the code. Each looked like its own bug. All five were one mechanism, and the fix is three greps.

The Wall We Had Already Published

A finding was published on this site, imported by a second project that saved hours with it, and rediscovered the hard way by a third five weeks later. The handbook stated it in the imperative three days before we shipped the bug. The knowledge was not missing — the structure that would have made us consult it was.

How to Fail Legibly

A literature-mining engine was asked, under a frozen bar, whether it could surface drug–disease connections that were validated only after its corpus ended. It couldn't — twice, under two configurations. The useful part is what made that answer worth keeping: a control that could have produced a positive, outcomes written before the data, three causes located by three instruments, and a closure rule that refuses to fire on an unlocated null.