Pith. sign in

REVIEW 5 major objections 6 minor 36 references

TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking

T0 review · 5 major / 6 minor · reviewed 2026-08-02 · deepseek-v4-flash

Pith's one-line read A language model trained only on 1800–1875 texts is structurally unable to know the future, and this paper argues that its resulting hallucinations are not errors but readable acts of nineteenth-century sensemaking.

desk verdict Original proposal and honest qualitative work, but the 'cannot know' claim rides on an unverified corpus-purity assumption and the quantitative evidence is thinner than it looks. read the letter →

arxiv 2607.24750 v1 pith:A25RKRCK submitted 2026-05-22 cs.CL cs.HC

classification cs.CLcs.HC
keywords selectivetemporaltrainingepistemologicaleventhorizongenerativearchivehallucinationasinterpretationVictorianliteraturehistoricallanguagemodelscomputationalhermeneuticsontologicalrepair
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper presents a 1.2-billion-parameter language model trained from scratch on roughly 16 billion words of Victorian prose, 1800–1875, with the claim that this temporal cutoff does more than shape style: it makes the model unable to represent post-1875 concepts. Because the model cannot 'pretend' to ignore the airplane or the computer—it simply has no representation of them—its wrong answers become reconstructions built from Victorian semantic resources. The authors call this 'ontological repair' and treat hallucination as a method for historical sensemaking. Alongside this, they report that two humanities scholars misclassified a large share of genuine Victorian excerpts as machine-written, which they read as a broader crisis of authenticity in how provenance is judged. If the core claim holds, deliberately bounded training data is a design choice that converts ignorance into an interpretive instrument.

What carries the argument

The load-bearing mechanism is 'selective temporal training': building a causal language model from scratch on a corpus whose publication window ends in 1875, including a tokenizer trained on that corpus, so that the model's only representational primitives are Victorian. The 'epistemological event horizon' is the cutoff date beyond which no world-knowledge can enter; it converts hallucinations into 'ontological repair', inferences drawn exclusively from pre-1875 semantic resources. A second piece is the use of held-out perplexity and vector projections (for example, projecting 'TIME' onto a NATURE–FACTORY axis) to argue that the model internalizes era-specific semantic structure.

What would settle it

Find a concrete passage in the training corpus or in generated outputs that uses a post-1875 concept in its modern sense—say, 'telephone' as an electric voice-transmission device or 'computer' as a calculating machine—and the temporal barrier is leaky. A direct test would be to prompt the model with 'telephone' or 'internet' and check whether any response reflects knowledge unavailable in 1875, or to scan the corpus for reprints of texts first published after 1875.

Watch

Extended reading notes

Core claim

On its own terms, the paper's central claim is that genuinely historical generation requires not prompt-level suppression of modern knowledge but structural exclusion of it. By training a 1.2-billion-parameter transformer-style causal model solely on digitized texts published between 1800 and 1875, the authors create an 'epistemological event horizon' at 1875; everything after is terra incognita. The model's perplexity on held-out Victorian prose falls 45.4% relative to a standard baseline, and its latent space shows 'TIME' aligning more strongly with 'FACTORY' than in a modern model. When prompted with anachronistic terms like 'computer', it produces historically grounded analogies—the comp

Load-bearing premise

Everything about the model's 'cannot know' status rests on the corpus containing only texts actually composed before 1876; the paper's check (that no post-1875 years appear in the text) would not catch modern reprints, editorial introductions, or OCR of later editions that carry post-1875 concepts without explicit dates.

Editorial extensions

If this is right

  • If the claim is right, historical simulation should be built with era-specific small models rather than prompted large models, because prompting cannot remove post-1875 weight structure.
  • Hallucination metrics for historical models should be reinterpreted: outputs that are factually wrong about the present may be right as evidence about the semantic field that produced them.
  • Archival honesty—not correcting historical biases—becomes a methodological requirement for studying empire and gender computationally; corrected models would obscure the structures under study.
  • The finding that experts misclassify real Victorian prose as machine-generated implies that provenance judgments are becoming less reliable, and that 'authentic' style may now be reproducible by design.
  • Temporal isolation may be measurable: the 45.4% perplexity improvement on period prose gives a concrete target for future temporally bounded models.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the corpus-purity assumption holds, the same method could produce bounded models for other periods or regions, allowing controlled comparisons of semantic drift across event horizons (for example, 1914 or 1945) rather than relying on one 1875 cutoff.
  • A testable extension would be probing the model with non-English or non-print concepts; the corpus is Anglocentric and print-based, so its 'Victorian ontology' is really the ontology of the imperial printed record, and the method's transferability to oral or colonial archives remains open.
  • The crisis of authenticity may generalize beyond this model: if a small 1.2-billion-parameter model can make Dickens look machine-written, then larger period-trained models might further erode textual provenance cues, with implications for digital archives and forgery detection.
  • The paper's reading of 'hypertrophied lung' as ontological repair could be tested quantitatively by asking whether the model's anachronism responses are consistently drawn from a small set of source domains (medicine, actuarial, maritime) and whether those domains track the most frequent collocations in the corpus.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 6 minor

Summary. The paper presents TimeCapsule, a 1.2B-parameter LLaMA-style causal model trained from scratch on an 89.63GB corpus of Internet Archive texts dated 1800–1875. The authors argue that this temporal restriction creates an 'epistemological event horizon' such that the model is structurally ignorant of post-1875 concepts; consequently, its hallucinations on anachronistic prompts can be read as 'ontological repair' that reveals Victorian sensemaking. Evidence includes a reported 45.4% perplexity reduction over GPT-2 on held-out Victorian prose, a diachronic semantic analysis showing 'TIME' shifting 2.1x toward 'FACTORY' relative to BERT, and a qualitative expert probe (N=2) in which experts misclassified 40–50% of genuine Victorian excerpts as machine-generated.

Significance. If the central epistemic claim could be established, TimeCapsule would be a valuable design paradigm: selective temporal training offers a concrete method for building historically constrained generative systems, and the reframing of hallucination as hermeneutic signal is thought-provoking. The paper is also transparent in releasing code, tokenizer, and checkpoints, and reports a carbon audit. However, the load-bearing assertion of temporal isolation is not yet empirically verified, and several quantitative comparisons are confounded. The contribution is potentially strong, but the current verification is insufficient.

major comments (5)
  1. [§3.2 / §2.1] The temporal-purity audit in §3.2 checks only that no post-1875 year strings appear in the corpus. This does not rule out post-1875 language entering through modern reprints, editorial introductions, critical apparatus, or OCR of later editions, all of which can be present in a digitized object dated 1800–1875 by Internet Archive metadata. Since the claim that the model 'cannot know' post-1875 concepts (§2.1) and the interpretations in §4.2 depend on the absence of modern semantic content, the authors need a content-level audit (e.g., frequency of post-1875 terms, provenance sampling, or a held-out classification task). Without this, the central isolation claim is unsupported. The limitations section (§6.4) acknowledges OCR noise but not temporal contamination.
  2. [§3.7.1] The headline perplexity comparison is not an apples-to-apples test of temporal isolation. GPT-2 is smaller and has a different tokenizer; Mistral-7B achieves lower raw PPL (16.50) and is dismissed without a temporal-isolation metric. To support the claim that 'period-specific linguistic fit' is improved by STT, the authors should compare models with the same architecture and tokenizer trained on historical vs modern or mixed corpora, and report a metric of post-1875 knowledge leakage (e.g., accuracy on anachronistic probes). As reported, the 45.4% reduction may reflect domain-specific training rather than the benefit of temporal restriction.
  3. [§4.1 / Figure 2] The diachronic semantic analysis rests on a single projection of 'TIME' onto a NATURE–FACTORY axis for TimeCapsule vs BERT. There are no error bars, no significance test, and no control for architecture (causal LM vs encoder), tokenizer, corpus scale, or domain. The 2.1x shift could be an artifact of any of these confounds. The authors should bootstrap the projections over multiple seeds and include an isomorphic baseline trained on a modern corpus under identical hyperparameters, with confidence intervals. The additional shifts ('VALUE→COMMERCE', 'POWER→STEAM') are reported without numbers or variance.
  4. [§4.2 / Table 3] The ontological-repair examples are cherry-picked: three hand-selected generations serve both to define and to demonstrate 'ontological repair'. The traced pathway (computer→calculation→vital statistics→lungs) is constructed after the fact from a single output. This is circular. A systematic protocol is needed: generate multiple completions per anachronistic prompt across different prompts and seeds, sample outputs blindly, and have domain experts judge historical plausibility and grounding without knowing the model. Failure cases (anachronistic or incoherent responses) should also be reported.
  5. [§5 / Figure 4] The expert probe is presented with quantitative framing (percentages, Cohen's kappa = 0.271) but rests on two experts and 40 classifications. The paper itself frames the study as a qualitative 'thick description' (§3.7.3), and in that capacity the observations are interesting. However, statements such as 'both experts misclassified approximately 40%' and 'crisis of authenticity' are not statistically supported. If the authors wish to retain the quantitative presentation, they need a larger panel and more items, or they should explicitly label these as illustrative observations from a small pilot.
minor comments (6)
  1. [§3.4] Tokenizer fertility reduction of 8.3–10.3% is reported without a definition of 'fertility' or a table of results. Please specify the exact metric and the models/tokenizers compared.
  2. [§3.7.2] The method for constructing semantic axes (e.g., averaged embeddings of definitional words, difference vectors) is underspecified. Reproducibility requires a precise description of how the NATURE–FACTORY axis is computed.
  3. [§4.1] Figure 2 caption says '7.5 (2.1x shift)' while the text reports BERT=6.67 and TimeCapsule=14.15. The ratio is 2.12 and the difference is 7.48; clarify which quantity '2.1x' refers to and whether the 7.5 is a projection score difference or ratio.
  4. [§5] The generated example contains stray '*' and '**' markers; the paper attributes them to OCR noise, but it does not report the decoding settings (temperature, top-p, repetition penalty). Please state these, as they affect output quality and reproducibility.
  5. [§6.2] The claim that 'This constraint cannot be reliably achieved through prompting or fine-tuning' is plausible but untested. A small control experiment with a modern model under explicit anti-anachronism prompting would strengthen this point.
  6. [References] Reference [35] lists 'Karl E Weick and Karl E Weick'; the author should appear once. Also fix the missing spaces in the abstract ('We presentTimeCapsule', 'TimeCapsuleexhibits').

Circularity Check

1 steps flagged · score 4.0 of 10

The central 'hallucination as ontological repair' claim is defined into existence: the same 'hypertrophied lung' output is used both to define and to demonstrate the phenomenon, while the quantitative PPL, embedding, and expert-probe results remain externally checkable.

  1. self definitional [§1 (definition), §2.6 (formal definition), §4.2 / Table 3 (demonstration); Abstract]
    "Instead, it performs an act of ontological repair: it generates interpretations grounded entirely in nineteenth-century semantic resources, for instance, imagining a computer not as a machine, but as a 'hypertrophied lung' (derived from actuarial tables of the period). ... The resulting description of a 'hypertrophied lung' is not simply a hallucination; it represents a logically constrained inference given the model's nineteenth-century epistemic horizon."

    The paper defines 'ontological repair' as the behavior of generating historically grounded interpretations when prompted with anachronisms, and the 'hypertrophied lung' output is introduced in §1 as the illustrative instance of that definition. In §4.2, the same output is then cited as evidence that the model 'represents a logically constrained inference' and that hallucinations are 'interpretive probes.' There is no independent criterion for what would count as a Victorian-sensemaking response versus an ordinary hallucination; the category and the evidence are the same token. Thus the central conclusion 'structural ignorance ... transforms hallucinations into interpretive probes' is entailed by the paper's own stipulation rather than established by measurement.

full rationale

TimeCapsule contains genuinely non-circular empirical content: a 1.2B model is trained from scratch on a declared 1800–1875 corpus; held-out perplexity (37.59 vs GPT-2 68.83, Mistral 16.50), tokenizer fragmentation, embedding projections, and the two-expert blind probe are all externally checkable measurements. There is no self-citation chain and no fitted parameter renamed as a prediction. The circularity is concentrated in the paper's central hermeneutic frame: 'ontological repair' is stipulated in §2.6 as the relational assembly of anachronistic prompts from Victorian primitives, and the same 'hypertrophied lung' example that defines the term in §1 is later offered in §4.2 as evidence that the model performs ontological repair. That is a definitional self-fulfillment rather than an empirical test. The temporal-purity concern—§3.2 only checks for absence of post-1875 year strings, not post-1875 concepts or contamination in digitized objects—is a real correctness risk to the 'cannot know' premise, but it is not itself a circularity, so it does not by itself raise the score. Overall the paper earns a moderate score: the quantitative evaluations provide independent grounding, but the most interesting interpretive claim is circularly constructed.

Assumptions & free parameters 4 free parameters · 5 assumptions · 3 invented entities

The central claims rest on several hand-picked design choices (cutoff year, semantic anchors, early stopping, tokenizer size) and on domain assumptions about what perplexity, embedding projections, and t-SNE clusters mean. The 'invented entities' are conceptual constructs rather than physical ones, but they do real work in framing the evidence.

free parameters (4)
  • Temporal cutoff year (epistemological event horizon) = 1875
    Chosen by the authors to precede telephone, electric light, and automobile; all historical-isolation claims depend on this boundary (§3.1).
  • Semantic axis anchors for diachronic analysis = 'NATURE' vs 'FACTORY' (and target 'TIME')
    Selected to test a pre-existing historical narrative; the reported 2.1× shift is relative to these hand-picked anchors (§4.1, Figure 2).
  • Early-stopping token budget = ≈13B tokens (50% of one epoch)
    Stopped when validation PPL stabilized; this model-selection choice is manual and affects the reported PPL (§3.5).
  • BPE vocabulary size = 32K
    Chosen to give single-token coverage to archaic spellings; tokenizer mismatch with GPT-2 complicates the headline PPL comparison (§3.4).
assumptions (5)
  • domain assumption Absence of post-1875 tokens in training data entails the model cannot represent or generate post-1875 concepts.
    Core premise of the epistemological event horizon; asserted in §2.1 and §3.2, but models can compose novel concepts from pre-1875 primitives and prompts can leak modern frames.
  • domain assumption Perplexity on held-out period prose is a meaningful measure of historical authenticity.
    Used in §3.7.1; the paper concedes lower raw PPL by Mistral-7B does not indicate authenticity, weakening PPL as evidence.
  • domain assumption Embedding projections from different model families (TimeCapsule vs BERT) are commensurable.
    §4.1 compares projection scores across a causal LLaMA-style model and a bidirectional BERT with different tokenizers; no alignment or normalization shown.
  • domain assumption Preserving historical biases is necessary for historical study (archival honesty).
    Normative stance in §2.4/§4.3; if false, the bias-topography evidence is not a strength but a harm.
  • domain assumption t-SNE visual clusters correspond to meaningful ideological structure.
    The bias topography in §4.3 interprets 2D t-SNE neighborhoods as evidence of colonial/gender ideologies; t-SNE preserves local structure approximately and can reflect token frequency and corpus composition.
invented entities (3)
  • epistemological event horizon
    purpose: A conceptual boundary at the training-data cutoff beyond which the model is claimed to have no world-knowledge; used to justify hallucinations as historical interpretation (§1, Figure 1).
    No falsifiable handle; it is a metaphor operationalized as the training-data cutoff.
  • latent materiality
    purpose: Treats model weights/embeddings as physical residues of historical discourse, legitimizing vector distances as historical evidence (§2.2).
    Philosophical extension of Kirschenbaum's forensic materiality; not independently measurable.
  • ontological repair
    purpose: The hypothetical process by which the model reconstructs anachronisms from Victorian primitives (§2.6, §4.2).
    Attributed to the model's behavior via selected examples; no mechanism-level evidence.

how reviews work

0 comments
Cite this review

Pith. "Pith review of TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking." pith.science (2026). https://pith.science/paper/A25RKRCK

@misc{pith2026260724750,
  author       = {Pith},
  title        = {Pith review of: TimeCapsule: Generative Hallucination as a Method for Historical Sensemaking},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/A25RKRCK}},
  note         = {Machine review of arXiv:2607.24750}
}
read the original abstract

Large Language Models (LLMs) are temporally overexposed: trained on vast contemporary corpora, they encode present-day concepts that make them unreliable narrators of the past. We present TimeCapsule, a 1.2B-parameter LLaMA-style causal model trained exclusively on Victorian texts (1800-1875) as an epistemologically isolated generative archive. Quantitative evaluation shows a 45.4% perplexity reduction over a GPT-2 baseline on held-out Victorian prose, while larger contemporary causal models achieve lower raw perplexity through broader pretraining but lack temporal isolation. TimeCapsule exhibits computational sensemaking, generating historically plausible analogical explanations for unfamiliar modern concepts (e.g., describing a computer as a "hypertrophied lung"). A qualitative hermeneutic probe with two humanities scholars revealed a crisis of authenticity, as both misclassified approximately 40% of genuine Victorian excerpts as machine-produced. We argue that structural ignorance of the future transforms hallucinations into interpretive probes of nineteenth-century ontologies.

Figures

Figures reproduced from arXiv: 2607.24750 by the authors.

Figure 1
Figure 1. The chronological boundary. The training data ter [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. The Semantic Shift. The projection of “TIME” [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. Bias Topography. In TimeCapsule, “progress” clus￾ters with “empire” and “dominion”. In Modern BERT, it clus￾ters with “improvement” and “machine”. These geometric patterns are further reflected in the model’s generative tendencies. The fact that TimeCapsule explicitly links progress to domination is not a defect but an instance of archival honesty. Rather than obscuring the “civilizing mission” logic of the British … view at source ↗
Figures from the paper (1 more)
Figure 4
Figure 4. Figure 4: Confusion matrix of expert judgments across 40 [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

36 extracted references · 1 linked inside Pith

  1. [1]

    2026.Sure, AI can ‘do’ writing

    Richard Beard. 2026.Sure, AI can ‘do’ writing. But memoir? Not so much. Aeon Magazine. https://aeon.co/essays/sure-ai-can-do-writing-but-memoir-not-so- much Accessed: 2026-01-31

  2. [2]

    2010.Georges Perec: A life in words

    David Bellos. 2010.Georges Perec: A life in words. Random House

  3. [3]

    Emily M Bender, Timnit Gebru, Angelina McMillan-Major, and Shmargaret Shmitchell. 2021. On the dangers of stochastic parrots: Can language models be too big?. InProceedings of the 2021 ACM conference on fairness, accountability, and transparency. 610–623

  4. [4]

    Jesse Josua Benjamin, Arne Berger, Nick Merrill, and James Pierce. 2021. Machine learning uncertainty as a design material: A post-phenomenological inquiry. In Proceedings of the 2021 CHI conference on human factors in computing systems. 1–14

  5. [5]

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. 2019. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers). 4171–4186

  6. [6]

    2020.Data Feminism

    Catherine D’Ignazio and Lauren F Klein. 2020.Data Feminism. MIT Press

  7. [7]

    2017.The interpretation of cultures

    Clifford Geertz. 2017.The interpretation of cultures. Basic books

  8. [8]

    Lars Hallnäs and Johan Redström. 2001. Slow technology–designing for reflection. Personal and ubiquitous computing5, 3 (2001), 201–212

Show all 36 references
  1. [9]

    William L Hamilton, Jure Leskovec, and Dan Jurafsky. 2016. Diachronic word embeddings reveal statistical laws of semantic change. InProceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 1489–1501

  2. [10]

    N Katherine Hayles. 2000. How we became posthuman: Virtual bodies in cyber- netics, literature, and informatics

  3. [11]

    Ryan Heuser. 2025. Generative Aesthetics: On formal stuckness in AI verse. (2025)

  4. [12]

    Nanna Inie, Jeanette Falk, and Raghavendra Selvan. 2025. How CO2stly is CHI? The carbon footprint of generative AI in HCI research and what we should do about it. InProceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–29

  5. [13]

    Nanna Inie, Peter Zukerman, and Emily M Bender. 2026. De-anthropomorphizing “AI”: From wishful mnemonics to accurate nomenclature.First Monday(2026)

  6. [14]

    2012.Mechanisms: New media and the forensic imagi- nation

    Matthew G Kirschenbaum. 2012.Mechanisms: New media and the forensic imagi- nation. mit Press

  7. [15]

    Gary Klein, Jennifer K Phillips, Erica L Rall, and Deborah A Peluso. 2007. A data–frame theory of sensemaking. InExpertise out of context. Psychology Press, 118–160

  8. [16]

    Cody Kommers, Ruth Ahnert, Maria Antoniak, Emmanouil Benetos, Steve Ben- ford, Mercedes Bunz, Baptiste Caramiaux, Shauna Concannon, Martin Disley, James Dobson, et al. 2025. Computational Hermeneutics: Evaluating Generative AI as a Cultural Technology. (2025)

  9. [17]

    Lucian Leahu. 2016. Ontological surprises: A relational perspective on machine learning. InProceedings of the 2016 ACM conference on designing interactive systems. 182–186

  10. [18]

    Makayla Lewis. 2025. Art, Identity, and AI: Navigating Authenticity in Creative Practice. InProceedings of the 2025 Conference on Creativity and Cognition. 916– 930

  11. [19]

    2020.Cultural analytics

    Lev Manovich. 2020.Cultural analytics. Mit Press

  12. [20]

    Jean-Baptiste Michel, Yuan Kui Shen, Aviva Presser Aiden, Adrian Veres, Matthew K Gray, Google Books Team, Joseph P Pickett, Dale Hoiberg, Dan Clancy, Peter Norvig, et al. 2011. Quantitative analysis of culture using millions of digitized books.science331, 6014 (2011), 176–182

  13. [21]

    2013.Distant reading

    Franco Moretti. 2013.Distant reading. Verso Books

  14. [22]

    William Odom, Ishac Bertran, Garnet Hertz, Henry Lin, Amy Yo Sue Chen, Matt Harkness, and Ron Wakkary. 2019. Unpacking the thinking and making behind a slow technology research product with slow game. InProceedings of the 2019 Conference on Creativity and Cognition. 15–28

  15. [23]

    William Odom and Tijs Duel. 2018. On the design of OLO Radio: Investigating metadata as a design material. InProceedings of the 2018 CHI Conference on Human Factors in Computing Systems. 1–9

  16. [24]

    William Odom, MinYoung Yoo, Henry WJ Lin, Tijs Duel, Tal Amram, and Amy Yo Sue Chen. 2020. Exploring the Reflective Potentialities of Personal Data with Different Temporal Modalities: A Field Study of Olo Radio. 283–295 pages

  17. [25]

    Fabian Offert and Peter Bell. 2020. Generative Digital Humanities.. InCHR. 202–212

  18. [26]

    2013.What is media archaeology?John Wiley & Sons

    Jussi Parikka. 2013.What is media archaeology?John Wiley & Sons

  19. [27]

    1969.La Disparition

    Georges Perec. 1969.La Disparition. Denoël

  20. [28]

    2018.New digital worlds: Postcolonial digital humanities in theory, praxis, and pedagogy

    Roopika Risam. 2018.New digital worlds: Postcolonial digital humanities in theory, praxis, and pedagogy. Northwestern University Press

  21. [29]

    Sangho Suh, Bryan Min, Srishti Palani, and Haijun Xia. 2023. Sensecape: En- abling multilevel exploration and sensemaking with large language models. In Proceedings of the 36th annual ACM symposium on user interface software and technology. 1–18

  22. [30]

    Mirac Suzgun, Tayfun Gur, Federico Bianchi, Daniel E Ho, Thomas Icard, Dan Jurafsky, and James Zou. 2025. Language models cannot reliably distinguish belief from knowledge and fact.Nature Machine Intelligence(2025), 1–11

  23. [31]

    EP Thompson. 1967. Time, Work-Discipline, and Industrial Capitalism.Past & Present38 (1967), 56–97

  24. [32]

    2019.Distant horizons: digital evidence and literary change

    Ted Underwood. 2019.Distant horizons: digital evidence and literary change. University of Chicago Press

  25. [33]

    Ted Underwood, Laura K Nelson, and Matthew Wilkens. 2025. Can Language Models Represent the Past without Anachronism?arXiv preprint arXiv:2505.00030 (2025)

  26. [34]

    Environmental Protection Agency

    U.S. Environmental Protection Agency. 2025. Emissions & Generation Resource Integrated Database (eGRID), Year 2023 Data. https://www.epa.gov/egrid. Re- leased January 2025

  27. [35]

    1995.Sensemaking in organizations

    Karl E Weick and Karl E Weick. 1995.Sensemaking in organizations. Vol. 3. Sage publications Thousand Oaks, CA

  28. [36]

    David Zhou, John R Gallagher, and Sarah Sterman. 2025. Thoughtful, Confused, or Untrustworthy: How Text Presentation Influences Perceptions of AI Writing Tools. InProceedings of the 2025 Conference on Creativity and Cognition. 573–589

Pith tools

Reviewed August 2, 2026 · model on record in the stance chip above.