What do language models know about the past? Recovering perspectival knowledge using large-scale language models
Date:
Poster Presentation at Text as Data 2026, Berkeley, CA, USA
Co-authors: Yash Mali and Laura Nelson
Abstract:
What, if anything do language models know about the past? While historians, and historical sociologists, are confined to the archive, langugage models, training on vast amounts of text, including historical text, may be able to generate, or synthesize, diffuse historical information, such as the prestige according to various occupations, when that information does not exist in the archives in a known form. We explore what types of historical knowledge is contained in base large language models (LLMs) in two stages. First, we validate token probabilities from prompted base models against census data, showing that some models can recover historical U.S. occupational distributions, once presentist bias is corrected via temporally-anchored prompts. Second, we attempt to recover undocumented occupational prestige hierarchies. Probing the same base model probability distributions over occupations recovers rankings loosely matching historical classifications, and the distributions shift predictably when prompts imply different perspectives (e.g., gendered or regional framing). We find that base models outperform instruction-tuned ones, probability-based probing outperforms generated text, and perspective can be deliberately shifted through prompting. We offer a preliminary pipeline for asking whether LLMs can simulate historical collective mentalities that are otherwise inaccessible to historical and comparative research.
