Pith. sign in

REVIEW 3 major objections 6 minor

Capability Provenance in Language Models: A Case Study in Social Reasoning

T0 review · 3 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read Social and STEM reasoning in a language model draw support from different, nameable regions of the pretraining corpus — and the split is sharper for reasoning than for factual knowledge.

desk verdict A genuine new instrument for capability provenance—but the sharper-for-reasoning claim is confounded by query-count imbalance, and the validation holds for SocialIQA only. read the letter →

arxiv 2606.19625 v3 pith:XSEYM7GX submitted 2026-06-17 cs.CL cs.LG

classification cs.CLcs.LG
keywords training-dataattributioncapabilityprovenancesocialreasoningmachineunlearningpretrainingcorpustopic-formattaxonomybenchmarkcontrastopen-datalanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish where a language model's capabilities come from in the pretraining corpus — specifically, whether social reasoning and STEM reasoning are fed by different, nameable regions of the training data. It maps gradient-based attribution scores onto a 576-cell topic-by-format taxonomy of the Dolma3 corpus and finds SocialIQA is the structural outlier: its support profile barely correlates with the three comparison benchmarks (r ≤ 0.22 vs 0.53–0.86 among them), and its strongest positive region, Literature × Customer Support, is strongly negative for the knowledge and STEM benchmarks. The social–STEM split is sharper for reasoning than for knowledge, and a single corpus bin flips sign between reasoning and knowledge in a correctness analysis. Targeted machine unlearning gives partial causal backing: forgetting high-attribution bins degrades the aligned benchmark more than random in-topic forgetting, clearly for SocialIQA (BH-adjusted p ≈ 1e-5), and this selectivity — though not the exact map — replicates on two other open-data models. If right, the paper supplies an inspectable map of which kind of text feeds which capability, with direct uses in data curation and capability auditing.

What carries the argument

The load-bearing unit is the WebOrganizer bin: one topic × format cell in a 24-topic × 24-format taxonomy (576 named regions, e.g. Literature × Customer Support). Noisy document-level gradient attribution (TrackStar influence on base-model gradients, queried through the instruction-tuned checkpoint) is aggregated into these bins, and each benchmark's profile is z-scored across all 576 cells so profiles can be compared by position rather than magnitude. The 2×2 benchmark design — domain (social vs STEM) crossed with capability type (reasoning vs knowledge) — supplies the contrastive frame, and NGDiff LoRA unlearning on the base model supplies the interventionist check of whether high-attribut

What would settle it

Pretrain or continue-pretrain a same-scale model with the top-attributed SocialIQA documents (concentrated in Literature × Customer Support) held out of the mix, alongside a matched random holdout; if the benchmark drop does not reproduce, the attribution map is not causal. A cheaper check: restrict the SocialIQA forget set to top-attributed documents drawn outside the Literature topic — if the Wilcoxon effect vanishes, the result is topic membership rather than attribution ranking.

Watch

Extended reading notes

Core claim

The central claim: in OLMo3-7B, social and STEM reasoning draw on qualitatively different corpus regions, and the contrast is stronger for the reasoning pair (SocialIQA vs ARC-Challenge) than for the knowledge pair (MMLU Social Sciences vs MMLU STEM). SocialIQA's 576-bin influence profile correlates with the three comparison benchmarks at r ≤ 0.22, while those three correlate at r = 0.53–0.86 among themselves. The most discriminative bin, Literature × Customer Support, is extreme-positive for SocialIQA (+16.0 z) and strongly negative for MMLU Social Sciences (−7.31) and MMLU STEM (−5.75); it is also the top correctness-differential cell for both reasoning benchmarks and the strongest error c

Load-bearing premise

The causal-validation leg assumes that how much unlearning damage a document causes tracks how strongly that document's gradient aligns with the benchmark's queries; the paper's own results show this holds clearly for SocialIQA but is null for MMLU Social Sciences, has a negative paired difference for ARC-Challenge, and does not co-rank for held-out Theory-of-Mind probes — if unlearning damage does not index the attribution signal, only the descriptive provenance profile rema

Editorial extensions

If this is right

  • Social reasoning in OLMo3-7B is carried heavily by short, dialogue-rich narrative text — the Literature × Customer Support bin is its single strongest support region and is simultaneously a negative influence on knowledge and STEM benchmarks.
  • Corpus support separates by capability type, not only domain: the social–STEM split is sharper between SocialIQA and ARC-Challenge than between their knowledge counterparts.
  • The same corpus content can help correct reasoning and hurt correct factual recall: Literature × Customer Support is the top correctness bin for both reasoning benchmarks and the strongest error bin for both knowledge benchmarks.
  • Forgetting high-attribution bins produces a selective damage profile rather than generic degradation, and this selectivity — though not the specific bin map — replicates across three open-data model families.
  • High-influence bins carry a consistent lexical signature — short interpersonal text alongside long-form documentation — so capability-relevant corpus regions are recognizable from cheap surface text features.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The Literature × Customer Support result suggests a genre-level hypothesis the paper leaves implicit: Q&A-style talk about characters and stories may train the model to track minds, not to store facts. A testable extension is a small-scale pretraining run that up-weights character-dialogue Q&A and checks held-out Theory-of-Mind gains.
  • The paper's G.8 finding — topical unlearning works for Theory-of-Mind while gradient attribution and unlearning damage do not co-rank — hints at a gap between gradient-aligned (typical) support and causally necessary (critical) data; separating the two could be its own measurement task.
  • The ARC-Challenge and MMLU Social Sciences validation failures suggest the causal leg is strongest where influence concentrates in a few unusual bins; for benchmarks with diffuse support, document-level ranking within a broad topic may be the wrong granularity, and a coarser or hierarchical bin structure could be tested.
  • If provenance maps are ecosystem-specific while causal selectivity transfers, social reasoning looks less like stored social knowledge and more like a reading style — implying curators could steer it with generic narrative data rather than domain-specific documents.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper develops a capability-provenance method: gradient-based training-data attribution (TrackStar via Bergson) on OLMo3-7B/Dolma3, aggregated over WebOrganizer's 24x24 topic-format taxonomy (576 bins), and applied in a 2x2 design crossing domain (social vs. STEM) with capability type (reasoning vs. knowledge): SocialIQA, MMLU Social Sciences, ARC-Challenge, MMLU STEM. The central descriptive finding is that SocialIQA's 576-bin influence profile is a structural outlier, correlating at r <= 0.22 with the three comparison benchmarks versus 0.53-0.86 among them (permutation p ~ 1e-4), with Literature x Customer Support extreme-positive for SocialIQA (+16.0 z) and negative for both MMLU benchmarks, and with the social-STEM separation reported as sharper at the reasoning level (max topic |Delta| = 0.91) than at the knowledge level (0.63). NGDiff targeted unlearning of high-attribution documents within topics degrades SocialIQA more than random in-topic controls (paired Wilcoxon BH-adjusted p ~ 1e-5), with weaker or null results for the other benchmarks. Lexical profiling, correctness differentials, and held-out probes (BBQ, ToMBench, BBH, etc.) provide corroborating breadth. Code, sampling manifests, the bin-level influence matrix, and unlearning checkpoints are released.

Significance. If the descriptive findings survive robustness checks, this is a useful, well-scoped contribution: it converts noisy document-level attribution into stable, named corpus regions and produces specific, falsifiable findings (the Literature x Customer Support bin as a social-reasoning-supporting region; its correctness-differential sign flip between reasoning and knowledge benchmarks). The paper earns explicit credit for an open, reproducible pipeline (detailed compute accounting, sampling manifests, released influence matrix and checkpoints), a genuine 2x2 factorial design, and notably candid reporting: Section 4.3 and Section 5 state that unlearning validation is null for MMLU Social Sciences and negative for ARC-Challenge, and G.8 openly reports that attribution and unlearning damage do not co-rank for held-out ToM probes. The descriptive claim is supported by several independent aggregations (full-grid correlations, marginals, permutation test). The significance is tempered by two gaps: the RQ2 'sharper' claim is potentially confounded by benchmark query-count imbalance, and the causal pillar currently supports only SocialIQA among the four primary benchmarks. The paper's value is

major comments (3)
  1. [Section 3.3 (aggregation rule), Table 8, Table 11, Appendix D.1] RQ2's claim that the social-STEM split is sharper at the reasoning level compares SocialIQA-ARC-Challenge with MMLU SS-MMLU STEM. Per-bin means average influence over benchmark queries (Section 3.3); ARC-Challenge has 1,172 queries vs SocialIQA's 10,000, while the MMLU pair is nearly balanced (3,077/3,018) (Table 8). ARC's bin means therefore carry roughly sqrt(10000/1172) ~ 2.9x more query-sampling variance, which survives within-benchmark z-standardization and inflates |Delta z| for the reasoning pair. Table 11's max-topic metric (0.91 vs 0.63) is exactly the kind of max statistic that variance overdispersion inflates. Both bootstrap schemes (Section 3.3 document-resampling; D.1 bin-resampling) hold the query set fixed, so this variance is unbounded in the paper. Figure 15's pattern (ARC correlating 0.53/0.55 with the MMLUs vs 0.86 between the MMLUs) is consistent with ARC being noise-
  2. [Section 4.3, Appendix G.8, Section 5] The causal-validation pillar assumes NGDiff unlearning damage on OLMo3-7B Base tracks the Instruct-queried attribution signal. For the primary benchmarks, validation is significant only for SocialIQA; MMLU Social Sciences is null, ARC-Challenge has a negative median paired difference, and MMLU STEM is weak (Section 4.3). For held-out ToM probes, G.8 reports that gradient attribution and unlearning damage do not co-rank (topic-level Spearman ~ 0) - a direct failure of that premise on the only held-out causal test. The paper is candid about this (G.8, Section 5), but Section 4.3's RQ3 takeaway ('usually causes larger accuracy drops') and the abstract's 'partial causal validation' somewhat overstate coverage: as reported, the causal pillar supports SocialIQA specifically, with the premise contradicted elsewhere. Recommend: (i) bring the G.8 non-convergence into the main-text RQ3 discussion;
  3. [Abstract (arXiv metadata); full text; appendices] The abstract states: 'We validate on other open-data model, Comma v0.1 7B-2T (Common Pile) and DCLM-Baseline-7B (DataComp-LM): causal selectivity holds on both models, while the provenance map is ecosystem-specific.' I could not find these experiments anywhere in the full text, the appendices, the resource accounting, or the released-artifact description; the full-text abstract itself omits this sentence. A claim of cross-ecosystem causal selectivity is load-bearing for the abstract's credibility and is not verifiable from the manuscript. Either add the experiments (with the same detail as Appendix J) or remove/correct the claim.
minor comments (6)
  1. [Section 4.1 vs Appendix D.1 / Figure 15] Section 4.1 says SocialIQA's profile correlates with the three comparison benchmarks 'at r <= 0.22 versus r = 0.76-0.86 among them,' but Appendix D.1 and Figure 15 report the among-comparison correlations as 0.53, 0.55, 0.86. Align the main text with the appendix range.
  2. [Section 3.4.1 vs Section 4.3] Section 3.4.1 specifies forget-set size k = 2,000, while Section 4.3's headline result is described as 'the faithful top-200 per-document replication.' State which k underlies Table 49/Figure 42 and where the k = 2,000 runs are reported.
  3. [Figure 12 / Appendix A] Figure 12's axis labels abbreviate topic names ('Sci. & Tech.', 'Social', 'Software Dev.') inconsistently with Table 2's canonical names (Science & Technology, Social Life, Software Development); add a caption note or use the canonical names.
  4. [Section 2] The claim of being 'the first application of gradient-based corpus-scale TDA to capability provenance with taxonomic aggregation' is broad; soften or delimit it against the cited prior work (e.g., Ruis et al. 2025, Zhao et al. 2024).
  5. [Figure 2 / Table 6 / Appendix H] Literature x Customer Support has a reported z = +16.0 for SocialIQA. Because this bin is in the corpus's token tail (Table 6) and the working set draws 10,000 of its ~34K documents, show the within-bin influence distribution (or a document-downsampling check) to rule out a small set of outlier documents driving the extreme mean.
  6. [Section 4.3 vs Section 5] Section 4.3 reports p ~ 1e-5 (Wilcoxon BH-adjusted) while Section 5 reports pBH < 1e-4 for the SocialIQA validation; label the test, conditioning, and multiplicity correction identically so the reader can compare.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; only a minor, non-load-bearing self-citation of the Bergson toolchain.

full rationale

The paper's central claims are associational measurements, not derivations from fitted parameters. Per-document TrackStar/Bergson influence scores are aggregated over the 576 WebOrganizer bins using the averaging rule in §3.3, and the RQ1/RQ2 conclusions (SocialIQA profile outlier; sharper social/STEM separation for reasoning than knowledge) are computed directly from those fixed influence estimates against four external benchmarks. Nothing in the attribution model is fit to the benchmark outcomes to produce the profiles, so the descriptive contrasts are not circular by construction. The unlearning study is a real intervention on OLMo3-7B Base and is not rigged: §4.3 reports that MMLU Social Sciences does not reach significance and ARC-Challenge has a negative median paired difference, and Appendix G.8 explicitly states that for held-out ToM probes gradient attribution and unlearning damage do not co-rank, disclaiming any forced convergence. The one self-citation is Bergson (Quirke et al., 2026), maintained by co-authors, but the attribution algorithm cited is the external TrackStar method (Chang et al., 2025) and the code is released, so this is a toolchain detail rather than load-bearing self-citation. The query-count imbalance between SocialIQA (10,000) and ARC-Challenge (1,172) is a real statistical robustness concern for the RQ2 contrast, but it is a correctness/validity risk rather than a circularity: the reported quantities are not defined in terms of query counts. I find no step in the derivation that reduces to its own inputs.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central measurement rests on (a) projected gradient alignment standing in for true training influence — fragile per the paper's own citations; (b) Base-document/Instruct-query checkpoint mixing with an identity second-order approximation; (c) NGDiff unlearning on Base being a valid test of Instruct-queried attribution; (d) WebOrganizer classifier accuracy on 1.1B documents; and (e) stratified equal-weight sampling supporting claims about the natural corpus. The four free parameters are design constants, not fitted values, and none enters the result by construction. No invented entities.

free parameters (4)
  • Stratified per-bin sampling budget = 10,000 documents per bin (5,678,621 docs / 10.5B tokens total)
    Hand-chosen equal per-bin budget over 576 bins. It fixes the noise floor of bin-level influence means and makes the maps conditional (equal-weight) rather than proportional to the corpus (Sec B.3); the abstract's phrasing 'which regions of the pretraining corpus support' inherits this design choice.
  • Unlearning forget-set size k = 2,000 documents per topic
    Sec 3.4.1: both random and influence-guided conditions use k=2,000; the magnitude and power of the validation contrast depend on this scale.
  • TrackStar projection dimension = d=16 (256 per module)
    Appendix C.4: Rademacher projection dimension; the approximation quality of the gradient alignment depends on it.
  • LoRA rank for NGDiff unlearning = rank 8
    Sec 3.4: rank-8 LoRA adapters merged into Base; the selectivity of forgetting depends on adapter capacity.
assumptions (5)
  • domain assumption TrackStar projected-gradient alignment approximates the true influence of training documents on benchmark predictions
    Sec 3.2 / Appendix C.1: s(j,i) = sum over modules of projected gradient dot products. The paper itself cites Basu et al. 2021 and Grosse et al. 2023 on influence-function fragility; all downstream bin-level claims inherit this approximation.
  • domain assumption The supervised instruction-tuning second-order term is identity, so Base-document gradients can be compared with Instruct-query gradients
    Sec 3.2: 'under the approximation used by Ruis et al. (2025) that the supervised-instruction-tuning second-order term is the identity.' This makes attribution an Instruct-queried, Base-indexed hybrid.
  • domain assumption NGDiff/LoRA unlearning on OLMo3-7B Base selectively removes the influence of target documents, so accuracy loss gamma measures causal dependence of the benchmark on that corpus region
    Sec 3.4: unlearning is applied to Base while attribution is Instruct-queried; a global random control nets out generic damage, but selective removal is assumed. The ARC-Challenge negative result and G.8 non-co-ranking are evidence against this assumption in those settings.
  • domain assumption WebOrganizer topic/format classifiers assign documents to the 576 bins with accuracy sufficient for bin-level claims
    Sec 3.1 / Appendix B: classification of 1.1B documents with gte-base-en-v1.5 classifiers; accuracy is deferred to Appendix L.1 as 'reported accuracy' and not stated where binning is introduced.
  • domain assumption The stratified equal-weight working set supports inference about how the natural Dolma3 corpus supports capabilities
    Sec B.3: stratified sampling deliberately does not preserve corpus proportions and the paper frames it as a conditional question, yet the abstract claims a mapping over 'the pretraining corpus'. The extreme bin Literature x Customer Support is about 0.01B tokens (near-empty tail) yet drives the headline.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Capability Provenance in Language Models: A Case Study in Social Reasoning." pith.science (2026). https://pith.science/paper/XSEYM7GX

@misc{pith2026260619625,
  author       = {Pith},
  title        = {Pith review of: Capability Provenance in Language Models: A Case Study in Social Reasoning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XSEYM7GX}},
  note         = {Machine review of arXiv:2606.19625}
}
read the original abstract

We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We validate on other open-data model, Comma v0.1 7B-2T (Common Pile) and DCLM-Baseline-7B (DataComp-LM): causal selectivity holds on both models, while the provenance map is ecosystem-specific. We open-source all code, data artifacts, influence scores, and checkpoints at https://github.com/eilab-gt/capabilibara and https://huggingface.co/HCAI-Lab.

Figures

Figures reproduced from arXiv: 2606.19625 by the authors.

Figure 1
Figure 1. Capability provenance pipeline. 1) Dolma3 binned by WebOrganizer topic– format; 2) Bergson/TrackStar attributes benchmark probes to bins; 3) signed z-scores map supportive and suppressive bins; 4) unlearning high-influence bins versus random in-topic controls tests causal effects. 1 arXiv:2606.19625v1 [cs.CL] 17 Jun 2026 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. SocialIQA is the provenance-structure outlier, with a bin-level profile that cor￾relates with each comparison benchmark at only r≤0.22 versus r = 0.53–0.86 among the three (Appendix D.1). A–B show marginal mean signed influence (z) over the 576-bin grid by format and topic; Documentation drives the comparison benchmarks, whereas Customer Support, Literature, and Social Life are positive only for SocialIQA. C highlig… view at source ↗
Figure 3
Figure 3. Schematic of the WebOrganizer cross-product taxonomy. Each of 24 topic cate [PITH_FULL_IMAGE:figures/full_fig_p017_3.png] view at source ↗
Figures from the paper (45 more)
Figure 4
Figure 4. Figure 4: Marginal distribution by WebOrganizer topic in the de-duplicated Dolma3 corpus. [PITH_FULL_IMAGE:figures/full_fig_p020_4.png]
Figure 5
Figure 5. Figure 5: Marginal distribution by WebOrganizer format in the de-duplicated Dolma3 [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: Joint topic–format distribution across the 576 WebOrganizer bins. [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Stratified versus representative sampling across WebOrganizer bins (token counts, [PITH_FULL_IMAGE:figures/full_fig_p025_7.png]
Figure 8
Figure 8. Figure 8: Difference between stratified and representative sampling allocations ( [PITH_FULL_IMAGE:figures/full_fig_p026_8.png]
Figure 9
Figure 9. Figure 9: Empirical cumulative distribution (ECDF) of document word counts across the de [PITH_FULL_IMAGE:figures/full_fig_p027_9.png]
Figure 10
Figure 10. Figure 10: Joint topic–format heatmap of mean document word count on a log color scale. [PITH_FULL_IMAGE:figures/full_fig_p028_10.png]
Figure 11
Figure 11. Figure 11: Per-benchmark influence score distributions across all 576 WebOrganizer bins. [PITH_FULL_IMAGE:figures/full_fig_p029_11.png]
Figure 12
Figure 12. Figure 12: Signed mean influence (z-scores) by topic (rows) and format (columns). Top row: [PITH_FULL_IMAGE:figures/full_fig_p032_12.png]
Figure 13
Figure 13. Figure 13: Mean signed influence by WebOrganizer topic, aggregated across formats. Values [PITH_FULL_IMAGE:figures/full_fig_p033_13.png]
Figure 14
Figure 14. Figure 14: Mean signed influence by WebOrganizer format, aggregated across topics. Values [PITH_FULL_IMAGE:figures/full_fig_p034_14.png]
Figure 15
Figure 15. Figure 15: Between-benchmark profile divergence shows [PITH_FULL_IMAGE:figures/full_fig_p035_15.png]
Figure 16
Figure 16. Figure 16: Bin-level social-vs-STEM separation. Signed differences compare [PITH_FULL_IMAGE:figures/full_fig_p036_16.png]
Figure 17
Figure 17. Figure 17: Paired topic-level influence comparison: SocialIQA vs. ARC-Challenge. Each [PITH_FULL_IMAGE:figures/full_fig_p037_17.png]
Figure 18
Figure 18. Figure 18: Paired topic-level influence comparison: MMLU Social Sciences vs. MMLU STEM. [PITH_FULL_IMAGE:figures/full_fig_p038_18.png]
Figure 19
Figure 19. Figure 19: Correctness differential: mean influence on correctly answered queries minus [PITH_FULL_IMAGE:figures/full_fig_p039_19.png]
Figure 20
Figure 20. Figure 20: Signed TrackStar influence for PUB after excluding [PITH_FULL_IMAGE:figures/full_fig_p042_20.png]
Figure 21
Figure 21. Figure 21: Signed mean influence for BBH Snarks (held-out benchmark) on the canonical [PITH_FULL_IMAGE:figures/full_fig_p043_21.png]
Figure 22
Figure 22. Figure 22: Signed TrackStar influence for BBH Disambiguation QA on the canonical topic– [PITH_FULL_IMAGE:figures/full_fig_p044_22.png]
Figure 23
Figure 23. Figure 23: Signed TrackStar influence for ToMBench on the canonical topic–format grid. [PITH_FULL_IMAGE:figures/full_fig_p046_23.png]
Figure 24
Figure 24. Figure 24: Signed TrackStar influence for the NegotiationToM belief/desire mask. The panel [PITH_FULL_IMAGE:figures/full_fig_p047_24.png]
Figure 25
Figure 25. Figure 25: Signed TrackStar influence for the SimpleToM mental-state subset. We show the [PITH_FULL_IMAGE:figures/full_fig_p048_25.png]
Figure 26
Figure 26. Figure 26: Signed TrackStar influence for BBQ. BBQ is included both as a strong direct [PITH_FULL_IMAGE:figures/full_fig_p049_26.png]
Figure 27
Figure 27. Figure 27: Signed TrackStar influence for MMLU moral/humanities. This panel gives [PITH_FULL_IMAGE:figures/full_fig_p050_27.png]
Figure 28
Figure 28. Figure 28: Signed TrackStar influence for MORABLES. The panel is retained as secondary [PITH_FULL_IMAGE:figures/full_fig_p051_28.png]
Figure 29
Figure 29. Figure 29: Signed TrackStar influence for MoralExceptQA/RBQA. We report it as secondary [PITH_FULL_IMAGE:figures/full_fig_p052_29.png]
Figure 30
Figure 30. Figure 30: Correctness differential for BBQ: signed influence on correctly answered queries [PITH_FULL_IMAGE:figures/full_fig_p053_30.png]
Figure 31
Figure 31. Figure 31: Correctness differential for ToMBench: signed influence on correctly answered [PITH_FULL_IMAGE:figures/full_fig_p054_31.png]
Figure 32
Figure 32. Figure 32: Topic-marginal signed TrackStar influence for the main hold-out masks plus the [PITH_FULL_IMAGE:figures/full_fig_p055_32.png]
Figure 33
Figure 33. Figure 33: Cross-probe supportive consensus for the top topic-format bins in Table [PITH_FULL_IMAGE:figures/full_fig_p056_33.png]
Figure 34
Figure 34. Figure 34: Held-out Theory-of-Mind unlearning. Per-subtask net influence ( [PITH_FULL_IMAGE:figures/full_fig_p057_34.png]
Figure 35
Figure 35. Figure 35: Row-normalized lexical profiles for the top-20 bins of the four primary bench [PITH_FULL_IMAGE:figures/full_fig_p060_35.png]
Figure 36
Figure 36. Figure 36: Format-cluster composition of the top-20 high-influence bins for the scoped [PITH_FULL_IMAGE:figures/full_fig_p060_36.png]
Figure 37
Figure 37. Figure 37: Pooled target-group contrasts (∆ = mean group A − mean group B over top-20 bins), with 95% document-bootstrap confidence intervals. The legend defines the compact contrast codes: C–E is commonsense reasoning minus expository targets, S–E is social reasoning minus expo…
Figure 38
Figure 38. Figure 38: OpenLIWC-style summary-variable proxies (Analytic-open, Tone-open, Clout [PITH_FULL_IMAGE:figures/full_fig_p063_38.png]
Figure 39
Figure 39. Figure 39: Benchmark-level OpenLIWC-style family loadings. Each cell is the mean full [PITH_FULL_IMAGE:figures/full_fig_p064_39.png]
Figure 40
Figure 40. Figure 40: Grouped OpenLIWC-style family loadings over pooled top-20 high-influence [PITH_FULL_IMAGE:figures/full_fig_p065_40.png]
Figure 41
Figure 41. Figure 41: Dimension-level OpenLIWC-style loadings for the 18 realized LIWC-adjacent [PITH_FULL_IMAGE:figures/full_fig_p066_41.png]
Figure 42
Figure 42. Figure 42: Unlearning validation Targeted machine unlearning validates bin-level attribution. For each WebOrganizer topic and each primary benchmark, γinfluence (colored markers: in￾topic, top-200 documents by per-document influence on the target benchmark) is compared against γ…
Figure 43
Figure 43. Figure 43: Cross-benchmark γ for single-bin unlearning across all 24 WebOrganizer topics, with γ = Abaseline − Aunlearned. Positive values indicate accuracy damage; negative values indicate an accuracy increase after unlearning. The vertical divider separates the topic intervent…
Figure 44
Figure 44. Figure 44: Per-benchmark γ for each of the 24 topic bins, sorted independently within each benchmark panel. Panel titles and borders encode benchmark role; muted-brown/gray bars encode positive accuracy damage versus accuracy increases relative to the zero line. J.3 Global Rando…
Figure 45
Figure 45. Figure 45: compares the γ distribution from all 24 single-bin runs (boxplots) against the global random control (star) for each benchmark. The control provides a procedural baseline for interpreting the topic-specific interventions: effects near the control are difficult to sepa…
Figure 46
Figure 46. Figure 46: Topic-level unlearning efficiency gain relative to the in-topic random baseline. [PITH_FULL_IMAGE:figures/full_fig_p086_46.png]
Figure 47
Figure 47. Figure 47: Relationship between Bergson influence scores and raw accuracy damage across [PITH_FULL_IMAGE:figures/full_fig_p087_47.png]
Figure 48
Figure 48. Figure 48: On-target γ (base − unlearned accuracy) for ARC unlearning under three task definitions on the full 5.68M working set, mean ± std over 25 recipes. All three sit at the noise floor [PITH_FULL_IMAGE:figures/full_fig_p092_48.png]

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.