The author built an AI‑assisted research pipeline to mine digitised historical archives for previously unknown events. Starting with the GLOBALISE transcriptions of the Dutch East India Company (4.35 million pages, 1600‑1790) plus Dutch and American newspapers and ship logs, they converted each passage into a semantic fingerprint (≈5.7 million vectors) using a GPU‑overnight run. A lightweight “System One” filter model named Jev screened the vectors for relevance (e.g., identifying animal mentions) at a cost of a few cents per million words; only the passages Jev flagged were passed to the larger Claude Haiku model for detailed reading, translation, and entity extraction, while a Claude Code agent verified each candidate against the original scanned documents. This two‑stage approach reduced manual review cost and time compared with using Haiku alone. The pipeline was first validated by rediscovering known events such as Breen’s 1615 dodo eyewitness account, the 1783 Laki eruption, and the 1815 Tambora eruption. Applying it to new queries yielded: a meteorite fall reported in the 19 December 1812 Java Government Gazette (a 4‑lb iron‑rich stone from Maharashtra, predating the region’s earliest recorded fall by 26 years); three previously undocumented Javan rhinoceros shipments to the King of Kandy (1738‑1740) with sworn crew testimony, a named ship (Loverendaal), and mortality records; and three volcanic eruptions absent from the Smithsonian Global Volcanism Program—Gamkonora (Halmahera, 3 Feb 1722), Ciremai (West Java, early 1712), and Slamet (Central Java, Feb 1780)—each corroborated by contemporary letters and geographic cross‑checks. The workflow has been released as the open‑source toolkit Antiquity for reproducible historical AI research.
Read original
hackernews