{"id":"c4613383-67db-4507-b9b4-ee7c3a70d5be","arxiv_id":"2608.11171","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Truthfulness grew from zero papers in 2021-2022 to the largest topic by 2025-2026, while explainability declined and then resurged in 2026 through mechanistic interpretability.","lead":"This paper classifies all 144 papers published at the TrustNLP workshop from 2021 to 2026 into six trust dimensions and traces how research priorities shifted over time. It offers a longitudinal map of trustworthy NLP research, relevant to anyone tracking how the field responds to major AI capability milestones.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The field-wide generalization rests on a cross-venue comparison that never applies the same inclusion criteria to TrustNLP and the *CL venues and reports no statistical test; until that comparison is made valid, representativeness is asserted, not demonstrated.","rationale":"The paper's strongest claim is that TrustNLP's evolution tracks the broader NLP community's trust priorities. The only quantitative evidence for that generalization is Appendix C. Reading in good faith, the workshop corpus itself is a real contribution: 144 papers, longitudinal coverage, and human-plus-LLM annotation with reported agreement. But the field-wide inference fails if the comparison set is not comparable. The asymmetry in inclusion criteria is concrete: keyword-title-filtered *CL papers are compared to an unfiltered TrustNLP corpus, and the same Sonnet classifier labels both, so classifier priors could manufacture the apparent match. The absence of any statistical test or uncertainty measure means Figure 2's 'closely follows' cannot be evaluated. The arithmetic inconsistencies (2,169 candidates versus 2,169 post-exclusion classified papers; abstract's 37% versus Table 2's 26% and the 34% computation for 2025-2026) compound the concern by showing that the quantitative reporting in the central appendix is not yet reliable. These are correctable, so conditional acceptance remains the right verdict; the stress test identifies exactly why the representativeness condition cannot yet be considered met.","tokens_in":20320,"tokens_out":6089,"duration_ms":55710,"concrete_test":"Release the cross-venue data and rerun the comparison under symmetric inclusion: apply the same 30-keyword title filter to TrustNLP that was applied to ACL/NAACL/EACL/EMNLP, classify a random sample of 200 *CL papers with a second independent annotator or human to estimate LLM agreement, and report per-year, per-dimension proportions with bootstrap confidence intervals or a chi-square test of TrustNLP versus the field average. If TrustNLP still falls within the field distribution after symmetric filtering, the representativeness claim survives; if not, the conclusions should be reframed as workshop-specific.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Appendix C is the only quantitative support for the headline claim that TrustNLP mirrors the field, and it does not support that inference. The *CL comparison set is built by a title-keyword filter, while TrustNLP is entered as the full workshop corpus (or a subset of 136); the two sides are not subject to the same inclusion rule. Both sides are then classified by the same Claude Sonnet 5 prompt, so any systematic label bias is shared and cannot be detected; there is no human-validation sample for the *CL papers. Figure 2 shows aggregate proportions only — no per-year breakdown, no error bars, and no test of whether TrustNLP's distribution differs from the field average. Appendix C also reports '2,169 papers' as both the pre-exclusion candidate set and the post-exclusion classified set, which is arithmetically inconsistent as written. Relatedly, the abstract's '37%' for truthfulness in 2025–2026 contradicts Table 2's own totals (38/144 = 26%; 27/79 = 34%). These are fixable, but until the comparison is recomputed with identical filtering and proper uncertainty quantification, the paper's central generalization is asserted rather than established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper analyzes all 144 archival TrustNLP workshop papers from 2021 to 2026, classifying them along six trust dimensions (Fairness & Bias, Robustness & Adversarial, Privacy, Machine Ethics & Safety, Truthfulness, Explainability) using one human and two LLM annotators. It reports temporal trends, a phase-by-phase synthesis of technical contributions, four structural insights, and a cross-venue comparison with ACL, NAACL, EACL, and EMNLP intended to show that TrustNLP mirrors the broader field's trust priorities. The abstract frames the main finding as a field-wide shift from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems.","tokens_in":20574,"tokens_out":6106,"duration_ms":55288,"significance":"If the claims hold, the paper would provide a valuable longitudinal record of how a stable ACL-affiliated workshop responded to capability changes in LLMs, and its full annotation table (Appendix D) is a useful resource for future meta-analyses. The use of two independent LLM classifiers with human adjudication, with agreement statistics reported, is a strength. However, the central generalization that TrustNLP is representative of the broader NLP community rests on Appendix C, and that evidence is currently not valid as implemented, due to inconsistent inclusion criteria and missing statistical support. The internal trend analysis from Table 2 is self-contained and useful even if the representativeness claim is withdrawn.","major_comments":[{"comment":"The truthfulness figures are internally inconsistent. The abstract states that truthfulness comprises 37% of papers by 2025–2026, but Table 2 gives 14 (2025) + 13 (2026) = 27 papers out of 79 for those two years, which is 34.2%, while the body's §3.3 reports 38 total and 26% of the corpus. The 37% figure appears to match only the single year 2025 (14/38 = 36.8%). Please reconcile these numbers, since the claim that truthfulness is the fastest-growing dimension is a headline quantitative result.","section":"Abstract, §3.3, Table 2"},{"comment":"Figure 2 labels the TrustNLP sample as n=136, while the paper states it analyzes all 144 proceedings papers. The discrepancy of 8 papers is never explained. Either the figure is using a subset (for example, excluding papers that failed classification), or the corpus count is wrong. This must be clarified because the figure is the only quantitative bridge to the field-generalization claim.","section":"Figure 2, Appendix C"},{"comment":"The cross-venue comparison does not apply the same inclusion rule to TrustNLP and the *CL venues. The ACL/NAACL/EACL/EMNLP set is filtered by title keywords, while TrustNLP appears to be entered as the full workshop corpus, so the comparison conflates the effect of keyword filtering with the effect of venue selection. In addition, Appendix C reports 2,169 papers both as the pre-exclusion candidate set and as the post-exclusion classified set, and Figure 2 shows only aggregate proportions with no per-year breakdown, no error bars, and no statistical test. Because this appendix is the only quantitative support for the paper's central claim that TrustNLP mirrors the field, the comparison must be recomputed with identical filtering and accompanied by uncertainty quantification, or the claim must be narrowed.","section":"Appendix C"},{"comment":"There are numerical mismatches between the narrative and Table 2 for the 2026 edition: §4.5 says explainability research comprises 11 papers in 2026, while Table 2 reports 13 Explainability papers; §4.5 also says Machine Ethics & Safety reaches its highest count with 8 papers, while Table 2 lists 9. Please correct the prose or the table so that the trend claims are reproducible from the data.","section":"§4.5, Table 2"},{"comment":"The claim of a causal or quasi-causal 'reactive lag' between capability events and topic shifts is stated as a structural insight, but the paper provides no quantitative test of the lag, and §7 itself acknowledges that the regularity is observed across only four capability emergences and that its future validity is an empirical question. The language in Insight 1 and the conclusion should be softened to a documented correlation, or a statistical analysis of event-to-topic lags should be added.","section":"§5.1 Insight 1, §7"}],"minor_comments":[{"comment":"The caption contains a typo: 'T oronto' should be 'Toronto'.","section":"Figure 1 caption"},{"comment":"The reference to 'Bui and V on Der Wense' contains stray spaces; the author's name should be 'Katharina von der Wense'.","section":"References"},{"comment":"The description of explainability as following a 'U-shaped trajectory' is not well supported by Table 2, which shows 5, 4, 2, 3, 2, 13 across years. The 2026 value is a sharp single-year spike rather than a smooth U-shape, so the characterization should be qualified or supported with additional analysis.","section":"§3.3"},{"comment":"Several entries in Table 3 use an em dash in the human or model columns (for example, the 2022 'An Encoder Attribution Analysis' row has human label '—'). The paper should state explicitly whether a dash means 'not a primary dimension' or 'missing annotation', since this affects how the final label set in Table 2 was derived.","section":"Appendix D"},{"comment":"The sentence 'In this way, we expand the scope of each class by merging overlapping class labels from the two taxonomies' is vague; please give a concrete example of how a TrustLLM label and a DecodingTrust label were merged into one of the six final dimensions.","section":"§3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for a journal audience, and the longitudinal dataset plus full annotation appendix is a genuine contribution. The main reason for major revision, rather than rejection, is that the weaknesses in the cross-venue comparison are fixable: the authors can recompute both sides with identical filtering, report per-year distributions with statistical tests, and reconcile the n=136 versus n=144 discrepancy. If the representativeness claim cannot be repaired, the authors should clearly limit the paper's scope to the TrustNLP corpus itself. I also note that one author served as the human annotator; this is disclosed, but the limitations section could more explicitly acknowledge the potential for annotator bias given the organizers' stake in the workshop's narrative."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a genuinely useful longitudinal map of the TrustNLP workshop, with a new annotated dataset of all 144 papers and a plausible narrative about how trust topics track capability releases. The numbers, though, are sloppier than they should be, and the cross-venue comparison that carries the \"workshop mirrors the field\" claim does not currently support it. The paper is worth engaging, but it needs revision, not acceptance as-is.\n\nWhat's new: a complete classification of the TrustNLP proceedings across six dimensions (Fairness, Robustness, Privacy, Ethics/Safety, Truthfulness, Explainability), released in Appendix D with annotation instructions and prompts. That is a real resource. The phase structure (interpretability/bias 2021-22, generative pivot 2023, trade-offs 2024, agentic/multimodal 2025, mechanistic/safety 2026) is well-grounded in the cited papers. The qualitative insights are mostly sensible; Insight 3 (output-level audits miss internal behavior) is the most substantive and is supported by both in-corpus and external citations. The authors also deserve credit for flagging their own scope limits in Section 7.\n\nSoft spots, in order of severity. First, the cross-venue comparison in Appendix C is the only quantitative support for the headline claim that TrustNLP mirrors the field, and it doesn't hold up. The *CL set is filtered by title keywords while TrustNLP enters as the full corpus (or 136 of 144, the figure is inconsistent), so the two sides are not compared on equal terms. There is no per-year breakdown, no error bars, no statistical test. The abstract's \"37%\" for truthfulness in 2025-2026 also contradicts Table 2's own numbers (38/144 = 26% overall; 27/79 = 34% for those two years), and Figure 2's n=136 doesn't match the 144 papers analyzed. These are all fixable, but until the comparison is recomputed with identical inclusion rules and uncertainty quantification, the paper's central generalization is asserted, not shown.\n\nSecond, calling explainability \"U-shaped\" is an overstatement: the trajectory is a decline to 2 papers in 2025 and a single-year spike to 13 in 2026, not a gradual recovery. The mechanistic interpretability story is real, but \"U-shaped\" flatters the data. Third, the annotation relies on one human annotator who is also an organizer. The dual role is disclosed and the LLM agreement is reported, but the final labels on the contested dimensions (Machine Ethics & Safety, Truthfulness) are set by that annotator, so the risk of subtle selection bias is real and not addressed.\n\nWho benefits: anyone working on trustworthiness in NLP who wants a quick, well-organized picture of how the field's priorities shifted, plus the dataset itself. It is not a breakthrough, but it is a solid descriptive contribution. Send it to peer review with a request for a careful revision of Appendix C, the headline numbers, and the U-shaped language. The core material is worth refereeing.","headline":"A genuinely useful longitudinal map of TrustNLP with a new 144-paper dataset, but the numbers are sloppy and the cross-venue generalization is currently asserted, not demonstrated.","tokens_in":21144,"tokens_out":3032,"would_cite":true,"duration_ms":24480,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The TrustNLP workshop's six editions show a field moving from post-hoc interpretability to mechanistic verification and proactive control, with each capability launch triggering a lagged shift in trust research.","keywords":["trustworthy NLP","TrustNLP","trust dimensions","truthfulness","mechanistic interpretability","capability emergence","longitudinal analysis","LLM safety"],"falsifier":"Classify trust-related papers from a different set of high-profile AI venues outside the NLP-conference family over 2021–2026 with the same six-dimension scheme and compare the year-by-year proportions; if truthfulness does not rise from absent to a dominant share after late 2022, or if explainability fails to rebound through mechanistic methods around 2026, the claimed field-wide transition and its lag structure would not generalize beyond TrustNLP.","tokens_in":20116,"feed_emoji":"📈","tokens_out":8964,"duration_ms":73355,"temperature":0.7,"pith_summary":"Analyzing all 144 papers published across six editions of the TrustNLP workshop (2021–2026), the paper establishes that the field's trust agenda is set by capability events: the release of high-impact chat models, open-weight models, and agentic and multimodal systems each triggered a lagged shift in what trust research asks. Classifying every proceedings paper along six trust dimensions, it finds truthfulness grew from absent in 2021–2022 to 13–14 papers per year in 2025–2026, fairness is the most persistent theme, and explainability follows a U-shape, declining as post-hoc attribution lost relevance and resurging in 2026 through mechanistic interpretability. A cross-venue comparison of roughly 2,000 trust papers from major NLP conferences shows the workshop's topical distribution tracks the field average, supporting the paper's central claim that TrustNLP is a representative window onto the community. The six-year arc moves from interpretability to reliability to controllability, so trust research is best understood as a response to what models can newly do, not as an internally driven theoretical progression. Getting this right matters because it implies that improving trust in AI means anticipating the next capability change rather than only refining evaluations of current systems.","feed_headline":"Six years of TrustNLP: research agenda tracks model launches","feed_subtitle":"Truthfulness jumped from zero to dominance; explainability vanished, then returned as mechanistic interpretability.","key_machinery":"The carrying mechanism is a six-category trust taxonomy built from TrustLLM's dimensions, refined with DecodingTrust's finer-grained distinctions, plus Explainability added because the corpus demanded it. Papers are multi-labeled by three annotators—one human and two large language models following identical annotation instructions—with human adjudication on disagreements, yielding overall decision-level accuracy above 90 percent. This taxonomy is applied twice: once to all 144 TrustNLP papers and once, after a 30-keyword filter, to about 2,169 trust papers from four major NLP conferences, so the trends can be checked against a field-level baseline. The chronological synthesis organizes the same corpus into five phases—interpretability and bias, the generative pivot, trust as a trade-off problem, agentic and multimodal frontiers, and mechanistic trust and safety at scale—that map shifts in the dominant human–model interaction mode.","core_discovery":"The central claim is that TrustNLP's six-year record documents a field-wide transition from post-hoc interpretability of static models to mechanistic understanding and proactive control of generative systems, and that this transition was driven by capability emergence rather than by cumulative theoretical progress. The paper's quantitative evidence is a classification of all 144 archival proceedings papers into six trust dimensions—Fairness & Bias, Robustness & Adversarial, Privacy, Machine Ethics & Safety, Truthfulness, and Explainability—derived from TrustLLM and DecodingTrust. The counts show truthfulness appearing only after the generative pivot and reaching 38 total papers, fairness present in every edition with 30 total, and explainability falling from 9 papers in 2021–2022 to 2 in 2025 before rebounding to 13 in 2026 through mechanistic interpretability. The authors read these trends as a lagged response to capability events: chat models activated all dimensions at once in 2023, open-weight and frontier models brought trade-off research in 2024, and agentic and multimodal systems pushed safety alignment and mechanistic verification in 2025–2026. They further argue from a cross-venue comparison of about 2,000 papers that TrustNLP's distribution mirrors the broader NLP community's priorities.","pith_inferences":["A testable extension would be to apply the same six-dimension classification to the next two workshop editions and check whether the observed one-edition lag between capability releases and topical shifts persists, since the authors flag that the regularity is retrospective.","The limitations section implies a sharper check the authors did not run: classifying trust papers at non-NLP-specific AI venues over the same period would reveal whether robustness or privacy grows faster than truthfulness there, which would narrow the representativeness claim to NLP venues.","An implicit consequence of the outputs-are-not-internals insight is that regulatory or deployment audits relying only on behavioral tests will miss latent misalignment, so audit protocols may soon need representation-level checks as a standard component.","The observed benchmark saturation suggests that evaluation sets should be maintained adversarially and updated as models improve; otherwise static benchmarks risk becoming training data and losing their measurement signal."],"forward_implications":["The truthfulness surge means hallucination, factuality, and calibration are now leading trust problems across the field, not niche concerns.","Mechanistic interpretability—probing and editing model internals—will keep displacing post-hoc attribution as the standard way to answer why a model behaved as it did.","Past capability shocks each produced a lagged shift in trust topics, so the next major modeling change should be followed within about a year by a visible jump in corresponding trust papers.","Output-level evaluation alone will increasingly be seen as insufficient, making internal probes a first-class component of trustworthy-evaluation practice.","Without a unifying framework, each new capability will continue to generate disconnected trust research lines rather than cumulative progress."],"supporting_citations":[{"why":"Supplies the 2021 proceedings that define the opening interpretability-and-bias phase analyzed in the paper.","marker":"Pruksachatkun et al., 2021"},{"why":"Supplies the 2022 proceedings that document the continuation of interpretability and fairness themes.","marker":"Verma et al., 2022"},{"why":"Provides the 2023 edition whose simultaneous activation of robustness, privacy, and truthfulness anchors the generative pivot finding.","marker":"Ovalle et al., 2023"},{"why":"Provides the 2024 edition containing the fairness–explainability and performance–fairness trade-off studies.","marker":"Ovalle et al., 2024"},{"why":"Provides the 2025 edition that documents the agentic and multimodal trust frontiers.","marker":"Cao et al., 2025"},{"why":"Supplies the six trust dimensions used as the base taxonomy for all annotations.","marker":"Huang et al., 2024"},{"why":"Refines scope boundaries and contributes finer-grained robustness sub-dimensions.","marker":"Wang et al., 2023"},{"why":"Positions the workshop's diachronic synthesis against the broader survey of trustworthy LLM evaluation.","marker":"Liu et al., 2023"},{"why":"Cited as outside evidence that internal representations predict factuality without generating text, supporting the outputs-are-not-internals insight.","marker":"Gottesman and Geva, 2024"}],"fun_headline_variants":["TrustNLP: truthfulness soars, explainability rebounds","From post-hoc to mechanistic: TrustNLP's six-year shift","Six years, 144 papers: TrustNLP mirrors field trends","Model launches drive trust research: TrustNLP evidence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that TrustNLP's six editions faithfully represent the trust research priorities of the entire NLP community; if the workshop's organizers, submission pool, or venue placement attract a skewed slice of the field, the claimed field-wide transition is really only a workshop-level trend.","fun_headline_variants_meta":{"raw":{"variants":["TrustNLP: truthfulness soars, explainability rebounds","From post-hoc to mechanistic: TrustNLP's six-year shift","Six years, 144 papers: TrustNLP mirrors field trends","Model launches drive trust research: TrustNLP evidence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000376,"raw_usage":{"total_tokens":2048,"prompt_tokens":1034,"completion_tokens":1014,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":650,"completion_tokens_details":{"reasoning_tokens":943}},"tokens_in":650,"tokens_out":1014,"duration_ms":8232,"temperature":1.0,"reasoning_tokens":943,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T04:46:36.736319+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Classify trust-related papers from a different set of high-profile AI venues outside the NLP-conference family over 2021–2026 with the same six-dimension scheme and compare the year-by-year proportions; if truthfulness does not rise from absent to a dominant share after late 2022, or if explainability fails to rebound through mechanistic methods around 2026, the claimed field-wide transition and its lag structure would not generalize beyond TrustNLP.","supporting_citations":[],"review_version":1}