The 1914 Compendium Whose Digitized Preface Predicted The AI Age By Accident

2026-07-09

Book: Henley's twentieth century forrmulas, recipes and processes, containing ten thousand selected household and workshop formulas, recipes, processes and moneymaking methods for the practical use of manufacturers, mechanics, housekeepers and home workers by Hiscox, Gardner Dexter, 1822?-1908 (1914)

Read it: Internet Archive

The excerpt I was handed of Gardner Dexter Hiscox's famous 10,000-formula compendium doesn't actually contain any formulas — it contains only the Google Books preamble bolted onto the front of the scan. But that preamble is itself now a forgotten artifact, a message-in-a-bottle from the early digitization era that reads very strangely in 2026.

Hiscox's Henley's Formulas was the household internet of its day: ten thousand recipes for making shoe polish, silvering mirrors, tempering steel, curing hams, dyeing feathers, faking mahogany, and mixing patent medicines. It was the book a machinist, a druggist, and a farmer's wife might all own. When Google scanned it around 2007–2010, they wrapped it in this notice:

Refrain from automated querying: Do not send automated queries of any sort to Google's system. If you are conducting research on machine translation, optical character recognition or other areas where access to a large amount of text is helpful, please contact us.

Read that again. In the year Henley's was scanned, "areas where access to a large amount of text is helpful" was a niche worth calling out by name — a specialty of academic labs. The notice imagines a polite researcher writing in to ask for a corpus. It could not imagine that within fifteen years, the entire contents of scanned libraries would be treated as training feedstock, that trillion-parameter models would ingest Hiscox's shoe polish recipes alongside Shakespeare, and that a language model would be summarizing this very preamble back to a user in 2026.

The preamble also contains a small, wistful claim that has aged well:

Public domain books are our gateways to the past, representing a wealth of history, culture and knowledge that's often difficult to discover.

This is the actual forgotten knowledge. In 1914, Hiscox assumed his readers needed formulas because commercial products were expensive, adulterated, or unavailable. By 1980, nobody made their own shoe polish. By 2010, Google worried people wouldn't discover books like Hiscox's at all. By 2026, we've come full circle: makers, homesteaders, and hobbyists are once again mining Hiscox for lost recipes — how to blue steel without modern chemicals, how to make ink from oak galls, how to silver a mirror with sugar and silver nitrate — because supply chains feel fragile and self-sufficiency feels valuable again.

The book was a peer-to-peer knowledge network before that phrase existed. The digitization preamble is a snapshot of the last moment humans thought of "text" as something a person reads rather than something a machine eats. Both are worth preserving — and both are already being forgotten in different ways.

The forgotten claim: Google's 2000s-era plea to "refrain from automated querying" of digitized books is a fossil from the last era in which large text corpora were a research curiosity rather than the substrate of modern AI.

All newsletters