A developer has compiled a 160GB dataset consisting of 40 billion tokens from 1800–1875 English text from England and the United States. Following the successful training of a 500M parameter evaluation model on a 5B token sample, plans are underway to train a 2B parameter model on the full dataset. The evaluation model has also undergone fine-tuning using synthetic 1800s-style Q&A pairs.

Read original