{"id":"c4d5cae3-da34-4f69-9538-3b7122ebc0ff","arxiv_id":"2504.19209","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"DETM's test-set word prediction is robust to loss reweighting and prior recomputation, but prefers larger vocabularies and fewer time windows across five diachronic corpora.","lead":"This paper tests how different setup choices affect Dynamic Embedded Topic Models using five historical text collections. It finds that vocabulary size and time-window granularity matter most, while several other choices barely change performance.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The abstract claim that loss reweighting has no significant effect conflicts with Table 3 under the paper's own 2σ=0.03 threshold; transferring that ACL-only threshold to one-run comparisons on other corpora compounds the problem.","rationale":"The Reader's weakest assumption was that the ACL-derived 2σ=0.03 threshold transfers to single-run results on the other four corpora. That concern is real and load-bearing. My stress-test found an additional, more direct problem: even if the threshold transfer were valid, the paper's own Table 3 contradicts the abstract's blanket 'not significantly or consistently affected' wording for loss reweighting. This is an internal inconsistency, not merely a cross-corpus generalization risk. The paper remains a useful empirical map, and the authors are transparent about their single-corpus significance estimate and limitations, but the headline claim needs to be narrowed or supported by multi-seed evidence on at least one non-ACL corpus. This reinforces the Reader's CONDITIONAL verdict rather than moving it to acceptance or rejection; the paper should be conditionally accepted with revisions to the abstract and significance discussion.","tokens_in":6472,"tokens_out":5656,"duration_ms":58401,"concrete_test":"Run the reweighting comparison with 20 independent seeds on at least one non-ACL corpus (e.g., UN, which shows a 0.07 difference), using the same fixed data split and training protocol. Compute the Bessel-corrected 2σ interval for each condition and the distribution of False-vs-True differences. If the UN difference is not robust across seeds, the 'no consistent effect' claim is supported; if it persists, the abstract's 'not significantly affected' claim must be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim (Abstract; §4) states that performance 'is not significantly or consistently affected by several aspects,' and the Reader's summary includes loss reweighting among those aspects. Table 3 does not support this for reweighting: using the paper's own Bessel-corrected 2σ=0.03 estimate from §2.6, the False-to-True NLL changes are acl 7.06→7.10, greek 6.70→6.88, latin 7.16→7.37, un 7.08→7.01. Four of five corpora move by more than 0.03, and the text itself says 'Three out of the four corpora showing significant differences favor not reweighting the loss.' That is a significant effect in most corpora, though the direction is not perfectly uniform. If the abstract intends to exclude reweighting from 'several aspects,' it should say so; as written, the central claim is internally inconsistent.\n\nIndependently, §2.6's threshold was estimated from 20 seeded runs on ACL only, while every other corpus appears with a single run per configuration. Differences like 0.04–0.21 in Table 3 therefore cannot be confidently called significant, nor confidently dismissed, for greek, latin, scifi, and un; the threshold transfer is an unverified assumption. Because the central 'no consistent effect' recommendation is grounded in these classifications, the headline conclusion is not yet established beyond the ACL corpus.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents an empirical study of implementation choices for Dynamic Embedded Topic Models (DETM) across five diachronic corpora (ACL, ancient Greek, Latin, science fiction, and UN). The authors vary six factors: recomputation of time-window statistics, loss reweighting, topic count, window count, vocabulary size, and the ratio of mixture-walk to topic-walk deltas, measuring held-out test-set negative log-likelihood (NLL), with NPMI coherence reported in an appendix. They conclude that several implementation aspects do not significantly or consistently affect NLL, that DETM is stable under vocabulary scaling, that smaller window counts perform best, and that larger topic counts generally improve performance. The paper releases standardized datasets, experimental code, and a pip-installable library.","tokens_in":6812,"tokens_out":5336,"duration_ms":52263,"significance":"The paper's main strength is its multi-corpus, held-out evaluation of practical DETM choices, addressing questions that arise in applied and humanistic scholarship. The release of datasets, code, and a library is a valuable contribution to reproducibility. If the empirical claims are fully supported after revision, the paper would be a useful map for practitioners. However, the inferential basis for the 'not significant' language is currently limited, and the abstract overstates the reweighting results. I therefore view the contribution as promising but requiring substantial qualification.","major_comments":[{"comment":"The abstract's claim that performance is 'not significantly or consistently affected by several aspects' is internally inconsistent with Table 3 and the surrounding text. Using the paper's own 2σ=0.03 threshold from §2.6, loss reweighting changes NLL by more than 0.03 on four of five corpora (acl 7.06→7.10, greek 6.70→6.88, latin 7.16→7.37, un 7.08→7.01), and the text explicitly states that 'Three out of the four corpora showing significant differences favor not reweighting the loss.' Reweighting is therefore significant in most corpora, even if the direction is not uniform. The abstract and §4 should be revised to name reweighting as an exception or to soften the significance language to 'small but not consistently directional effects.' As written, the central claim is not supported.","section":"Abstract and §4, Table 3"},{"comment":"The only variance estimate in the paper is based on 20 seeded runs on the ACL corpus, yielding 2σ=0.03, but every other corpus is reported as a single run per configuration. Applying this ACL-derived threshold to greek, latin, scifi, and un assumes that run-to-run variability is identical across corpora, which is not shown and is questionable given the large differences in corpus size and language. Consequently, statements such as 'None of the corpora show a significant difference' (Table 2) and the classification of differences as significant or not in Tables 3–7 are not supported for four of the five corpora. The authors should either provide per-corpus variance estimates (e.g., multiple seeds for at least the key comparisons) or explicitly present the 2σ value as an ACL-only heuristic and avoid strong significance claims elsewhere.","section":"§2.6, Tables 2–7"},{"comment":"The claim that DETM is 'very stable when scaling vocabulary size' is also in tension with the reported numbers. On the UN corpus, NLL worsens from 7.10 at 20k to 7.17 at 80k, a change larger than the paper's own 2σ=0.03 threshold. While the other four corpora are indeed stable, the single-run nature of the UN measurement means this difference cannot simply be dismissed, and the blanket stability claim should be qualified with explicit attention to this outlier.","section":"§3, Table 6"}],"minor_comments":[{"comment":"The text says 'We format in bold the best performance in a table row unless it doesn't meet this threshold,' but none of the tables in the paper contain bold entries; either the formatting was omitted or the sentence should be removed.","section":"§2.6"},{"comment":"There is a typo: 'Leuven Databast of Ancient Books' should be 'Leuven Database of Ancient Books.'","section":"§2.1"},{"comment":"The phrase 'We have conveniently decided to focus on basic training and convergence' is informal and could be replaced with a neutral statement of what the paper does and does not claim about topic interpretability.","section":"§2.4"},{"comment":"The references to ANONYMIZED dataset, code, and library URLs must be replaced with actual repository identifiers before publication, and Table 6's header 'V ocab Size' contains a stray space.","section":"Tables and code availability"}],"recommendation":"major_revision","confidential_remarks":"The paper has a useful empirical contribution and good reproducibility infrastructure, but the abstract overstates the findings and the significance methodology is not yet sufficient for the claims made about non-ACL corpora. I recommend major revision rather than rejection because the issues are addressable through careful rewriting of the claims and either additional seed runs or explicit qualification of the significance language."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper you should know about: a systematic sweep of DETM implementation choices across five genuinely diverse diachronic corpora. The valuable findings are that recomputing window statistics makes no difference, vocabulary size can be scaled 5k to 80k with almost no NLL movement, and small window counts are consistently best. That is exactly the kind of practical guidance applied users need, and it is new as a comparative dataset spanning Greek, Latin, sci-fi, UN, and ACL.\n\nWhat it does well: held-out test NLL, random splits at the document level, a fixed default configuration while sweeping one axis, an NPMI appendix that honestly reports the weak NLL-coherence correlation, and a limitations section that does not oversell. The writing is clear and the experimental hygiene is above average for this kind of work. The authors also built a library, though in this anonymized version neither code nor data is inspectable.\n\nThe soft spot is real and load-bearing. The abstract says performance is \"not significantly or consistently affected by several aspects\" and implies loss reweighting is among them. Table 3 disagrees under the paper's own 2σ = 0.03 threshold: reweighting moves NLL by 0.04, 0.18, 0.21, 0.03, and 0.07 across the five corpora. The text in Section 3 even says \"three out of the four corpora showing significant differences favor not reweighting.\" So the central null claim is internally inconsistent as written. The authors should either exclude reweighting from that sentence or soften \"not significantly\" to \"inconsistently and with mixed direction.\"\n\nThe second soft spot is the significance threshold. It was estimated from 20 seeded runs on ACL only, then applied to single runs on the other four corpora. That transfer is an assumption, and for those corpora differences of 0.02–0.04 cannot be confidently classified as significant or not. The paper flags this in Section 2.6, but then proceeds as if the threshold applies everywhere. That undercuts the confidence of several \"no significant effect\" statements, though not the recompute or vocabulary-size findings, which look robust.\n\nBottom line: this is a competent empirical contribution that deserves a serious referee, not a desk reject. It will be most useful to digital humanists and applied topic-modeling users who need to know which knobs to turn. I would recommend it for peer review with a request to fix the abstract/Table 3 contradiction and to treat the cross-corpus significance claims more cautiously. I'd cite it if I worked on DETM; I'd probably leave it out of a reading group unless someone is actively using the model.","headline":"Useful empirical map of DETM hyperparameter sensitivity, but the abstract overstates the null result on loss reweighting, and the significance threshold is carried across corpora on thin ice.","tokens_in":7276,"tokens_out":1943,"would_cite":true,"duration_ms":21008,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper maps which implementation choices matter for the Dynamic Embedded Topic Model (DETM) and finds that recomputing window statistics and reweighting the loss have no significant or consistent effect on test-set negative…","keywords":["dynamic embedded topic model","diachronic corpora","negative log-likelihood","hyperparameter sensitivity","vocabulary scaling","temporal windows","topic models","semantic change"],"falsifier":"Run 20 seeded replicates of the same configuration on each of greek, latin, scifi, and un, and compute the corpus-specific 2σ interval for test NLL. If that interval is wider than the observed differences—for instance, if latin's roughly 0.21 reweighting gap falls inside seed noise—then the paper's conclusion that these choices have no significant or consistent effect would not hold for those corpora.","tokens_in":6313,"feed_emoji":"📊","tokens_out":6249,"duration_ms":57290,"temperature":0.7,"pith_summary":"The paper asks which implementation choices actually matter when applying the Dynamic Embedded Topic Model (DETM) to historical text. Using five diachronic corpora covering ancient Greek, Latin, early science fiction, ACL papers, and UN proceedings, it measures test-set negative log-likelihood under controlled sweeps of five choices: recomputing window statistics, reweighting the loss, delta ratios, window count, and vocabulary size. The central finding is that two choices that complicate deployment—recomputing time-window statistics and reweighting the loss with a global batch ratio—do not significantly or consistently affect performance, so they can be dropped. Vocabulary size can be scaled from 5,000 to 80,000 with almost no change in NLL, and smaller window counts perform best, though the model degrades gracefully with more windows. The paper positions these results as an initial map for applied scholars and for a future continuous-time variant.","feed_headline":"Topic-model tweaks barely move performance on five diachronic corpora","feed_subtitle":"Across five corpora, only window count and topic count move the needle; vocabulary scales to 80k.","key_machinery":"The central object is the Dynamic Embedded Topic Model (DETM), a neural topic model that combines word embeddings with a dynamic topic model: each time window has a topic-mixture prior and topic embeddings that evolve by random walks, and a variational inference network maps sub-documents to topic mixtures. The paper's instrument is controlled before-and-after comparison of test-set per-word negative log-likelihood, with a Bessel-corrected estimate of $2\\sigma = 0.03$ derived from 20 seeded ACL runs to set a significance threshold. The argument works by holding all other hyperparameters at their defaults and sweeping one factor per corpus, then checking whether differences clear the threshold.","core_discovery":"The paper establishes that DETM's test-set negative log-likelihood is not significantly or consistently affected by recomputing the mixture-prior summary statistics at validation or test time, nor by reweighting the reconstruction and KL terms by the ratio of training documents to batch size. Across five corpora, NLL differences from these choices fall near or below a $2\\sigma = 0.03$ threshold estimated from 20 seeded ACL runs, and where differences are larger they do not favor one option consistently. The model is stable when vocabulary size grows from 5,000 to 80,000 words, continues to improve with larger topic inventories up to 80–160 topics on several corpora, and performs best with 2–4 temporal windows while remaining robust to over-granular windows through an interpolation policy for empty windows. $\\Delta$ ratios show no interpretable pattern except that the highest ratio is consistently worst.","pith_inferences":["Because the paper measures NLL rather than human-rated topic quality, its recommendation to drop recomputation and reweighting concerns predictive fit; a practitioner whose goal is interpretable topics should still check that those changes do not alter the topics surfaced.","The stability across vocabulary sizes suggests DETM is a good fit for historical text with orthographic variation and pre-standardized spelling, provided the softmax memory bottleneck is addressed; the paper notes this direction but does not test it.","A cheap protocol for future applied work follows from the significance caveat: before comparing configurations on a new corpus, run a small set of seeded replicates to calibrate corpus-specific noise rather than reusing the ACL threshold."],"forward_implications":["Practitioners can turn off recomputation of window statistics and loss reweighting, eliminating the need to track global batch statistics during training and inference.","DETM's predictive performance is essentially flat as vocabulary size grows from 5,000 to 80,000, so the practical limit is memory in constructing normalized categorical topic distributions rather than model quality.","Fewer temporal windows (2 or 4) give the best test NLL, and the interpolation policy keeps over-granular windows from causing catastrophic drift, supporting the pursuit of a continuous-time variant.","Larger topic inventories continue to improve NLL on several corpora, with no single optimum across all five; some corpora benefit well beyond 80 topics.","Delta-ratio choices have no consistent interpretation except that a 9:1 ratio is consistently worst, so defaulting to equal deltas is a reasonable starting point."],"supporting_citations":[{"why":"Introduces the Dynamic Embedded Topic Model, the model whose implementation choices this paper sweeps.","marker":"Dieng et al. (2019)"},{"why":"Supplies the dynamic topic model structure with random walks over time that DETM extends.","marker":"Blei and Lafferty (2006b)"},{"why":"Provides the latent Dirichlet allocation foundation and the classic topic-modeling setup DETM builds on.","marker":"Blei et al. (2003)"},{"why":"Motivates the vocabulary-scaling question by showing how embedding-based topic models handle rare words and long-tail vocabulary.","marker":"Dieng et al. (2020)"},{"why":"Documents the pitfalls of treating NLL as a proxy for human interpretability, which the paper cites to justify focusing on predictive fit.","marker":"Hoyle et al. (2021)"},{"why":"Supplies the Word2Vec skip-gram embeddings trained for the sub-document representations used in the experiments.","marker":"Mikolov et al. (2013)"},{"why":"Defines the normalized pointwise mutual information coherence metric reported in the appendix.","marker":"Bouma (2009)"}],"fun_headline_variants":["Most DETM tweaks don't matter, only topics and windows do","DETM: only topic count and windows matter","Five corpora show DETM robust to most tuning choices","Vocabulary scales to 80k, but DETM ignores most tweaks","Keep DETM simple: pick topics and windows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The significance threshold used to decide whether differences matter is estimated from 20 seeded runs on the ACL corpus only and then applied to the other four corpora, each of which is tested with a single run per configuration.","fun_headline_variants_meta":{"raw":{"variants":["Most DETM tweaks don't matter, only topics and windows do","DETM: only topic count and windows matter","Five corpora show DETM robust to most tuning choices","Vocabulary scales to 80k, but DETM ignores most tweaks","Keep DETM simple: pick topics and windows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000599,"raw_usage":{"total_tokens":2733,"prompt_tokens":813,"completion_tokens":1920,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":429,"completion_tokens_details":{"reasoning_tokens":1832}},"tokens_in":429,"tokens_out":1920,"duration_ms":11472,"temperature":1.0,"reasoning_tokens":1832,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:57:47.052859+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run 20 seeded replicates of the same configuration on each of greek, latin, scifi, and un, and compute the corpus-specific 2σ interval for test NLL. If that interval is wider than the observed differences—for instance, if latin's roughly 0.21 reweighting gap falls inside seed noise—then the paper's conclusion that these choices have no significant or consistent effect would not hold for those corpora.","supporting_citations":[],"review_version":1}