REVIEW 3 major objections 6 minor
Capability Provenance in Language Models: A Case Study in Social Reasoning
T0 review · 3 major / 6 minor · reviewed 2026-08-04 · deepseek-v4-flash
Pith's one-line read Social and STEM reasoning in a language model draw support from different, nameable regions of the pretraining corpus — and the split is sharper for reasoning than for factual knowledge.
desk verdict A genuine new instrument for capability provenance—but the sharper-for-reasoning claim is confounded by query-count imbalance, and the validation holds for SocialIQA only. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing unit is the WebOrganizer bin: one topic × format cell in a 24-topic × 24-format taxonomy (576 named regions, e.g. Literature × Customer Support). Noisy document-level gradient attribution (TrackStar influence on base-model gradients, queried through the instruction-tuned checkpoint) is aggregated into these bins, and each benchmark's profile is z-scored across all 576 cells so profiles can be compared by position rather than magnitude. The 2×2 benchmark design — domain (social vs STEM) crossed with capability type (reasoning vs knowledge) — supplies the contrastive frame, and NGDiff LoRA unlearning on the base model supplies the interventionist check of whether high-attribut
What would settle it
Pretrain or continue-pretrain a same-scale model with the top-attributed SocialIQA documents (concentrated in Literature × Customer Support) held out of the mix, alongside a matched random holdout; if the benchmark drop does not reproduce, the attribution map is not causal. A cheaper check: restrict the SocialIQA forget set to top-attributed documents drawn outside the Literature topic — if the Wilcoxon effect vanishes, the result is topic membership rather than attribution ranking.
Extended reading notes
Core claim
The central claim: in OLMo3-7B, social and STEM reasoning draw on qualitatively different corpus regions, and the contrast is stronger for the reasoning pair (SocialIQA vs ARC-Challenge) than for the knowledge pair (MMLU Social Sciences vs MMLU STEM). SocialIQA's 576-bin influence profile correlates with the three comparison benchmarks at r ≤ 0.22, while those three correlate at r = 0.53–0.86 among themselves. The most discriminative bin, Literature × Customer Support, is extreme-positive for SocialIQA (+16.0 z) and strongly negative for MMLU Social Sciences (−7.31) and MMLU STEM (−5.75); it is also the top correctness-differential cell for both reasoning benchmarks and the strongest error c
Load-bearing premise
The causal-validation leg assumes that how much unlearning damage a document causes tracks how strongly that document's gradient aligns with the benchmark's queries; the paper's own results show this holds clearly for SocialIQA but is null for MMLU Social Sciences, has a negative paired difference for ARC-Challenge, and does not co-rank for held-out Theory-of-Mind probes — if unlearning damage does not index the attribution signal, only the descriptive provenance profile rema
Editorial extensions
If this is right
- Social reasoning in OLMo3-7B is carried heavily by short, dialogue-rich narrative text — the Literature × Customer Support bin is its single strongest support region and is simultaneously a negative influence on knowledge and STEM benchmarks.
- Corpus support separates by capability type, not only domain: the social–STEM split is sharper between SocialIQA and ARC-Challenge than between their knowledge counterparts.
- The same corpus content can help correct reasoning and hurt correct factual recall: Literature × Customer Support is the top correctness bin for both reasoning benchmarks and the strongest error bin for both knowledge benchmarks.
- Forgetting high-attribution bins produces a selective damage profile rather than generic degradation, and this selectivity — though not the specific bin map — replicates across three open-data model families.
- High-influence bins carry a consistent lexical signature — short interpersonal text alongside long-form documentation — so capability-relevant corpus regions are recognizable from cheap surface text features.
Reading between the lines
- The Literature × Customer Support result suggests a genre-level hypothesis the paper leaves implicit: Q&A-style talk about characters and stories may train the model to track minds, not to store facts. A testable extension is a small-scale pretraining run that up-weights character-dialogue Q&A and checks held-out Theory-of-Mind gains.
- The paper's G.8 finding — topical unlearning works for Theory-of-Mind while gradient attribution and unlearning damage do not co-rank — hints at a gap between gradient-aligned (typical) support and causally necessary (critical) data; separating the two could be its own measurement task.
- The ARC-Challenge and MMLU Social Sciences validation failures suggest the causal leg is strongest where influence concentrates in a few unusual bins; for benchmarks with diffuse support, document-level ranking within a broad topic may be the wrong granularity, and a coarser or hierarchical bin structure could be tested.
- If provenance maps are ecosystem-specific while causal selectivity transfers, social reasoning looks less like stored social knowledge and more like a reading style — implying curators could steer it with generic narrative data rather than domain-specific documents.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops a capability-provenance method: gradient-based training-data attribution (TrackStar via Bergson) on OLMo3-7B/Dolma3, aggregated over WebOrganizer's 24x24 topic-format taxonomy (576 bins), and applied in a 2x2 design crossing domain (social vs. STEM) with capability type (reasoning vs. knowledge): SocialIQA, MMLU Social Sciences, ARC-Challenge, MMLU STEM. The central descriptive finding is that SocialIQA's 576-bin influence profile is a structural outlier, correlating at r <= 0.22 with the three comparison benchmarks versus 0.53-0.86 among them (permutation p ~ 1e-4), with Literature x Customer Support extreme-positive for SocialIQA (+16.0 z) and negative for both MMLU benchmarks, and with the social-STEM separation reported as sharper at the reasoning level (max topic |Delta| = 0.91) than at the knowledge level (0.63). NGDiff targeted unlearning of high-attribution documents within topics degrades SocialIQA more than random in-topic controls (paired Wilcoxon BH-adjusted p ~ 1e-5), with weaker or null results for the other benchmarks. Lexical profiling, correctness differentials, and held-out probes (BBQ, ToMBench, BBH, etc.) provide corroborating breadth. Code, sampling manifests, the bin-level influence matrix, and unlearning checkpoints are released.
Significance. If the descriptive findings survive robustness checks, this is a useful, well-scoped contribution: it converts noisy document-level attribution into stable, named corpus regions and produces specific, falsifiable findings (the Literature x Customer Support bin as a social-reasoning-supporting region; its correctness-differential sign flip between reasoning and knowledge benchmarks). The paper earns explicit credit for an open, reproducible pipeline (detailed compute accounting, sampling manifests, released influence matrix and checkpoints), a genuine 2x2 factorial design, and notably candid reporting: Section 4.3 and Section 5 state that unlearning validation is null for MMLU Social Sciences and negative for ARC-Challenge, and G.8 openly reports that attribution and unlearning damage do not co-rank for held-out ToM probes. The descriptive claim is supported by several independent aggregations (full-grid correlations, marginals, permutation test). The significance is tempered by two gaps: the RQ2 'sharper' claim is potentially confounded by benchmark query-count imbalance, and the causal pillar currently supports only SocialIQA among the four primary benchmarks. The paper's value is
major comments (3)
- [Section 3.3 (aggregation rule), Table 8, Table 11, Appendix D.1] RQ2's claim that the social-STEM split is sharper at the reasoning level compares SocialIQA-ARC-Challenge with MMLU SS-MMLU STEM. Per-bin means average influence over benchmark queries (Section 3.3); ARC-Challenge has 1,172 queries vs SocialIQA's 10,000, while the MMLU pair is nearly balanced (3,077/3,018) (Table 8). ARC's bin means therefore carry roughly sqrt(10000/1172) ~ 2.9x more query-sampling variance, which survives within-benchmark z-standardization and inflates |Delta z| for the reasoning pair. Table 11's max-topic metric (0.91 vs 0.63) is exactly the kind of max statistic that variance overdispersion inflates. Both bootstrap schemes (Section 3.3 document-resampling; D.1 bin-resampling) hold the query set fixed, so this variance is unbounded in the paper. Figure 15's pattern (ARC correlating 0.53/0.55 with the MMLUs vs 0.86 between the MMLUs) is consistent with ARC being noise-
- [Section 4.3, Appendix G.8, Section 5] The causal-validation pillar assumes NGDiff unlearning damage on OLMo3-7B Base tracks the Instruct-queried attribution signal. For the primary benchmarks, validation is significant only for SocialIQA; MMLU Social Sciences is null, ARC-Challenge has a negative median paired difference, and MMLU STEM is weak (Section 4.3). For held-out ToM probes, G.8 reports that gradient attribution and unlearning damage do not co-rank (topic-level Spearman ~ 0) - a direct failure of that premise on the only held-out causal test. The paper is candid about this (G.8, Section 5), but Section 4.3's RQ3 takeaway ('usually causes larger accuracy drops') and the abstract's 'partial causal validation' somewhat overstate coverage: as reported, the causal pillar supports SocialIQA specifically, with the premise contradicted elsewhere. Recommend: (i) bring the G.8 non-convergence into the main-text RQ3 discussion;
- [Abstract (arXiv metadata); full text; appendices] The abstract states: 'We validate on other open-data model, Comma v0.1 7B-2T (Common Pile) and DCLM-Baseline-7B (DataComp-LM): causal selectivity holds on both models, while the provenance map is ecosystem-specific.' I could not find these experiments anywhere in the full text, the appendices, the resource accounting, or the released-artifact description; the full-text abstract itself omits this sentence. A claim of cross-ecosystem causal selectivity is load-bearing for the abstract's credibility and is not verifiable from the manuscript. Either add the experiments (with the same detail as Appendix J) or remove/correct the claim.
minor comments (6)
- [Section 4.1 vs Appendix D.1 / Figure 15] Section 4.1 says SocialIQA's profile correlates with the three comparison benchmarks 'at r <= 0.22 versus r = 0.76-0.86 among them,' but Appendix D.1 and Figure 15 report the among-comparison correlations as 0.53, 0.55, 0.86. Align the main text with the appendix range.
- [Section 3.4.1 vs Section 4.3] Section 3.4.1 specifies forget-set size k = 2,000, while Section 4.3's headline result is described as 'the faithful top-200 per-document replication.' State which k underlies Table 49/Figure 42 and where the k = 2,000 runs are reported.
- [Figure 12 / Appendix A] Figure 12's axis labels abbreviate topic names ('Sci. & Tech.', 'Social', 'Software Dev.') inconsistently with Table 2's canonical names (Science & Technology, Social Life, Software Development); add a caption note or use the canonical names.
- [Section 2] The claim of being 'the first application of gradient-based corpus-scale TDA to capability provenance with taxonomic aggregation' is broad; soften or delimit it against the cited prior work (e.g., Ruis et al. 2025, Zhao et al. 2024).
- [Figure 2 / Table 6 / Appendix H] Literature x Customer Support has a reported z = +16.0 for SocialIQA. Because this bin is in the corpus's token tail (Table 6) and the working set draws 10,000 of its ~34K documents, show the within-bin influence distribution (or a document-downsampling check) to rule out a small set of outlier documents driving the extreme mean.
- [Section 4.3 vs Section 5] Section 4.3 reports p ~ 1e-5 (Wilcoxon BH-adjusted) while Section 5 reports pBH < 1e-4 for the SocialIQA validation; label the test, conditioning, and multiplicity correction identically so the reader can compare.
Circularity Check
No significant circularity; only a minor, non-load-bearing self-citation of the Bergson toolchain.
full rationale
The paper's central claims are associational measurements, not derivations from fitted parameters. Per-document TrackStar/Bergson influence scores are aggregated over the 576 WebOrganizer bins using the averaging rule in §3.3, and the RQ1/RQ2 conclusions (SocialIQA profile outlier; sharper social/STEM separation for reasoning than knowledge) are computed directly from those fixed influence estimates against four external benchmarks. Nothing in the attribution model is fit to the benchmark outcomes to produce the profiles, so the descriptive contrasts are not circular by construction. The unlearning study is a real intervention on OLMo3-7B Base and is not rigged: §4.3 reports that MMLU Social Sciences does not reach significance and ARC-Challenge has a negative median paired difference, and Appendix G.8 explicitly states that for held-out ToM probes gradient attribution and unlearning damage do not co-rank, disclaiming any forced convergence. The one self-citation is Bergson (Quirke et al., 2026), maintained by co-authors, but the attribution algorithm cited is the external TrackStar method (Chang et al., 2025) and the code is released, so this is a toolchain detail rather than load-bearing self-citation. The query-count imbalance between SocialIQA (10,000) and ARC-Challenge (1,172) is a real statistical robustness concern for the RQ2 contrast, but it is a correctness/validity risk rather than a circularity: the reported quantities are not defined in terms of query counts. I find no step in the derivation that reduces to its own inputs.
Assumptions & free parameters
free parameters (4)
- Stratified per-bin sampling budget =
10,000 documents per bin (5,678,621 docs / 10.5B tokens total)
- Unlearning forget-set size k =
2,000 documents per topic
- TrackStar projection dimension =
d=16 (256 per module)
- LoRA rank for NGDiff unlearning =
rank 8
assumptions (5)
- domain assumption TrackStar projected-gradient alignment approximates the true influence of training documents on benchmark predictions
- domain assumption The supervised instruction-tuning second-order term is identity, so Base-document gradients can be compared with Instruct-query gradients
- domain assumption NGDiff/LoRA unlearning on OLMo3-7B Base selectively removes the influence of target documents, so accuracy loss gamma measures causal dependence of the benchmark on that corpus region
- domain assumption WebOrganizer topic/format classifiers assign documents to the 576 bins with accuracy sufficient for bin-level claims
- domain assumption The stratified equal-weight working set supports inference about how the natural Dolma3 corpus supports capabilities
Cite this review
Pith. "Pith review of Capability Provenance in Language Models: A Case Study in Social Reasoning." pith.science (2026). https://pith.science/paper/XSEYM7GX
@misc{pith2026260619625,
author = {Pith},
title = {Pith review of: Capability Provenance in Language Models: A Case Study in Social Reasoning},
year = {2026},
howpublished = {\url{https://pith.science/paper/XSEYM7GX}},
note = {Machine review of arXiv:2606.19625}
}
read the original abstract
We use training-data attribution as an interpretable tool for capability discovery, mapping which regions of the pretraining corpus support social-reasoning versus STEM-reasoning in OLMo3-7B. Training-data attribution measures how strongly each training document influences a model's predictions on a benchmark, but document-level scores are too noisy to identify which corpus regions support which capabilities. We compute gradient-based attribution (TrackStar via Bergson) over a working set drawn from the de-duplicated Dolma3 mix, aggregate influence across WebOrganizer's 24-format x 24-topic taxonomy (576 bins), and contrast benchmark pairs in a 2x2 design that varies domain (social vs. STEM) and capability type (reasoning vs. knowledge): SocialIQA and MMLU Social Sciences against ARC-Challenge and MMLU STEM. Social and STEM reasoning draw on qualitatively distinct corpus regions, and the contrast is sharper at the reasoning level than at the knowledge level. Targeted machine unlearning provides partial causal validation: forgetting high-attribution topics (e.g., Literature for SocialIQA) degrades the aligned benchmark more than within-topic random baselines. We validate on other open-data model, Comma v0.1 7B-2T (Common Pile) and DCLM-Baseline-7B (DataComp-LM): causal selectivity holds on both models, while the provenance map is ecosystem-specific. We open-source all code, data artifacts, influence scores, and checkpoints at https://github.com/eilab-gt/capabilibara and https://huggingface.co/HCAI-Lab.
Figures
Figures from the paper (45 more)
Reviewed August 4, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.