Pith. sign in

Paper Citation Record · LEDGER

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2505.07293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.07293 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:23:36.717737Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T12:57:26.458109Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T06:07:41.034120Z

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy5
  • unresolved49
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 15052427-48b9-41a9-8b25-c77efce244c3 · outbound

This paper cites SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.342892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.342892Z digest=sha256:aaade0f717a68ca6e145502429c211ccc13801f6d765643cb2c8dcbca528c635

Observation 569a43c7-96b9-4512-8f85-cdd643740b12 · outbound

This paper cites Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Perplexed by Perplexity: Perplexity-Based Data Pruning With Small Reference Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.349667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.349667Z digest=sha256:acc97b78489ba15d996c8b5dc4e3533adafff6a91a7f51b1f46bebed66e45018

Observation b7ade84b-4b54-4018-a42e-bee1c430bb16 · outbound

This paper cites Smollm-corpus,.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Smollm-corpus,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.355299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.355299Z digest=sha256:f0d49c46dc9f73badc986d21005d4992f92ec2fa191a794a7654802bcbff173d

Observation 2c76a108-f35f-40b1-903f-9d77feff11d3 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Piqa: Reasoning about physical commonsense in natural language

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.368664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.368664Z digest=sha256:fed047050f65aa7560a9370ae8f366dee1ac6b63cccae5bcd9c5fdf58a4b13e6

Observation aed61845-245e-47aa-98eb-f335b73c0347 · outbound

This paper cites Towards monoseman- ticity: Decomposing language models with dictionary learning.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Towards monoseman- ticity: Decomposing language models with dictionary learning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.690423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:35.376604Z digest=sha256:456615dba62a89a2ec344325e224be2ba54f5f08fcc32b12bbe8d0bfda73c7a9

Observation 3cf91f5d-d6c8-42ce-925d-e88747df4240 · outbound

This paper cites Efficient Intent Detection with Dual Sentence Encoders.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Efficient Intent Detection with Dual Sentence Encoders

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.382927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.382927Z digest=sha256:f0501750c160e17cb0db8566e836eb601717749a97389777323c814a700264ae

Observation 1baea1ab-b71e-4cca-8a05-cd76f6d155e2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Evaluating Large Language Models Trained on Code

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.471433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.471433Z digest=sha256:2a052b652fba1e78f56d322d0a4990ac65294873d385785f8a6fd7ebc69f56e0

Observation c6c97e77-ca05-401c-85fe-9ce465d75f56 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.565205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.565205Z digest=sha256:afb1ea2d4b213a227affdb00bc2cd87bda7af7105d9deab9ae45ce4f9da25cc3

Observation c64a20fe-1d9e-4398-9eef-1b2ae8502e15 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.679161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.679161Z digest=sha256:742464f836b595cbfe4b5a522f28b588fcfa7216c3e3b9bb98852d740b4d22c7

Observation 4a00cb4c-a9d6-484b-9a20-5c363045ed9d · outbound

This paper cites DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DROP: A Reading Comprehension Benchmark Requiring Discrete Reasoning Over Paragraphs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.686363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.686363Z digest=sha256:1b074d86a5f0ca3199b3a1beed5d44d71872889a219be6a4a068916383e2fbab

Observation b36fbfd1-88ab-4571-8ea4-5434f4336146 · outbound

This paper cites Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Not all heads matter: A head-level kv cache compression method with integrated retrieval and reasoning.arXiv preprint arXiv:2410.19258, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.691752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.691752Z digest=sha256:cbc122a8600f949e7021926804829cead1aa65ded443d16fb01d8d6b15cb7fbe

Observation 6a40318b-0c9f-4437-9189-f6318105ea54 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Transformer Feed-Forward Layers Are Key-Value Memories

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.696873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.696873Z digest=sha256:95e0f40466b7272b0754eeecc47c860050fe112e3f8aac047152535f94275713

Observation 7acfe777-81cb-4bb6-a786-6e249ad97c33 · outbound

This paper cites The Llama 3 Herd of Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The Llama 3 Herd of Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.701654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.701654Z digest=sha256:5b8893c21a14c188c1665da3b1b2ceb92a57b0160c152b07e050ba1b16770fb4

Observation a88aba2a-ac26-4803-9af0-f023b5bd3126 · outbound

This paper cites Optimizing Pretraining Data Mixtures with LLM-Estimated Utility.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Optimizing Pretraining Data Mixtures with LLM-Estimated Utility

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.706278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.706278Z digest=sha256:ebf6768f3e1dd977a326271a37055f98056e3ed9bdcd6cff625c0a483982cf03

Observation 536a8b14-8458-4a31-87c7-0ef0a3cb05b5 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Measuring Massive Multitask Language Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.755937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.755937Z digest=sha256:ef343cb0578d2fa36c26cbb6afd70ac5ca778d8f3cd09d8d75216b4cf48d502d

Observation 2d03563a-249c-4f5d-b769-cf0bf78000ac · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.865394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.865394Z digest=sha256:7b4b904540fca2d97f42b3547edbecdf651c18448dfd1373b8e262c929d68ac1

Observation bdc98545-8281-4dcb-b2a6-0c9efa88a7f3 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Distilling the Knowledge in a Neural Network

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.972358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.972358Z digest=sha256:c5735efef1fcbdc1c913ff0d59247d6b7e9583705a917af9eeda17839c8c7982

Observation 5cbba9c2-fc87-44cc-907f-07727ec3d6d6 · outbound

This paper cites MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.976986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.976986Z digest=sha256:92905eb707b6cdd8406690da0e481da2181ec7e136eff37299f4cb2b1452683f

Observation 1306cc60-deb5-40fb-b1bd-9a2b326454ae · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.981579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.981579Z digest=sha256:2c39cba9831d29c6621b7070571139e34586ca0e355f8ad859c3d30183efc114

Observation 7624a982-d77e-4261-b3ba-4d94b28278ca · outbound

This paper cites TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.985894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.985894Z digest=sha256:dd1df25e5487b2f39b86acf14581d465130f83dcf25b83962ba9981445c140ea

Observation d18e8cd2-ecdc-4426-a0e5-248831889a75 · outbound

This paper cites FastText.zip: Compressing text classification models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection FastText.zip: Compressing text classification models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.991918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.991918Z digest=sha256:350f7a68e2e3812a3fc15d401db8b0653ec1946da4a9d98ce729ba1dae7bfaca

Observation 9bce7833-f631-49c4-98fb-4c6689a22566 · outbound

This paper cites The mirrored influence hypothesis: Efficient data influence estimation by harnessing forward passes.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The mirrored influence hypothesis: Efficient data influence estimation by harnessing forward passes

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.667541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:35.996508Z digest=sha256:915d06363a16be0840528967fb9726ea44491f791fffff156c6c9a85dd427eca

Observation 305aafe7-e1ab-470a-8712-979f547e8a11 · outbound

This paper cites RACE: Large-scale ReAding Comprehension Dataset From Examinations.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection RACE: Large-scale ReAding Comprehension Dataset From Examinations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.000905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.000905Z digest=sha256:de00c9a33a44d375ea71f29a150258d30b8fda7c821dbb4b34eef91d33f30578

Observation 5dbcc96b-45f7-4856-9eec-336deb053f64 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Datacomp-lm: In search of the next generation of training sets for language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.004896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.004896Z digest=sha256:7d90d85b1e8d14756964a1b556f2a6aefc246f23991247074cd495440e6a48d9

Observation 4bc5fe5b-f274-4d54-b097-8cebf7207d18 · outbound

This paper cites ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection ScalingFilter: Assessing Data Quality through Inverse Utilization of Scaling Laws

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.012927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.012927Z digest=sha256:8a2066c6a26588915d3feae6669ba4e25e5bd1994d12f85db486ebdc01593dee

Observation bf3bc0d5-060e-4834-b6ee-e68a1b4bf6a3 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Rho-1: Not All Tokens Are What You Need

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.017738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.017738Z digest=sha256:733e72a7f4b0b3ae2951b2fc1c4683fcdd1606e07353ccf1f11641cf3443d9a1

Observation 163f28b4-f51b-4165-a90c-3ea13af6f64c · outbound

This paper cites RegMix: Data Mixture as Regression for Language Model Pre-training.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection RegMix: Data Mixture as Regression for Language Model Pre-training

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.073298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.073298Z digest=sha256:a67fac6261b70ab35537205f6eb5406021e43ed8180a94edbee65021e853e628

Observation de5d1edb-5f33-423e-9b9f-477cb03c0a84 · outbound

This paper cites Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Interpreting Key Mechanisms of Factual Recall in Transformer-Based Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.181667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.181667Z digest=sha256:19e15998ce929da09e8ea41e08da492033c095b78408f519f75241de88383314

Observation d43b5d7a-c0cd-4cb6-b09f-f7d93bb999d9 · outbound

This paper cites Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.303495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.303495Z digest=sha256:8ef0725d63333e57890527c10dc57209e732c73b528f48934d672cdce1f3ac89

Observation d5fc0da0-896e-4e02-8ebf-acdffd5c36c4 · outbound

This paper cites 2 OLMo 2 Furious.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection 2 OLMo 2 Furious

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.309882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.309882Z digest=sha256:e47b11ab01a3e3f08675d17dab8d4c8b391dd8fcd0342325eeb08a6a36840f7f

Observation d7c0278b-4f75-47d2-bd93-13c13b5b2a1a · outbound

This paper cites In-context Learning and Induction Heads.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection In-context Learning and Induction Heads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.314555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.314555Z digest=sha256:68fd1dfd234c92288d5da94d231a118435986c01e7d66f1889d848bf2da733b9

Observation 549d502c-fe53-4c79-8ba9-790de5786f38 · outbound

This paper cites The fineweb datasets: Decanting the web for the finest text data at scale.Advancesin Neural Information Processing Systems, 37:30811–30849, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection The fineweb datasets: Decanting the web for the finest text data at scale.Advancesin Neural Information Processing Systems, 37:30811–30849, 2024

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.318551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.318551Z digest=sha256:4092821e2a6fa92915980c644d42106b56f05aacd358ada02910c66d0f150416

Observation 8e477ee4-2df2-4ae2-93ce-7c1326bd5a14 · outbound

This paper cites DataMan: Data Manager for Pre-training Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DataMan: Data Manager for Pre-training Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.323015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.323015Z digest=sha256:98ad36e454b715c57bc1d5fae030a3647900be0c5809c6cf0a204210f8bc963f

Observation eccd86e7-1eeb-4d3e-a5d7-efc6425c456b · outbound

This paper cites CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection CLongEval: A Chinese Benchmark for Evaluating Long-Context Large Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:23:37.109571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.326984Z digest=sha256:5507ae5f4650c52c032701aefb0772e2c2dc9447079df0e91a44579d3d30a850

Observation d9e3d4bc-8ec3-44a2-81f5-a5fe463174d9 · outbound

This paper cites Scaling Language Models: Methods, Analysis & Insights from Training Gopher.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Scaling Language Models: Methods, Analysis & Insights from Training Gopher

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.331205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.331205Z digest=sha256:036f0c3927acc39b3e2a9933864eab95a12f42890b4f456294276017e8d31daa

Observation b1a710f5-84fa-4872-b11f-228ddd6ecf9a · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.335197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.335197Z digest=sha256:19007e2d21055462469aca7f9a9bd947d28cd7393c14c98383fb401f6b3b2a2f

Observation e4f4725c-4c7c-4f97-a2d8-ab229eb7892a · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.433043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.433043Z digest=sha256:4988fea8ce6b8c286d4beac6358dc052efb82f0c32323484bfcbf7d69e16a724

Observation d00461b2-4989-4b1b-bdf6-2e1ae719bfb0 · outbound

This paper cites Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Procedural Knowledge in Pretraining Drives Reasoning in Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.485660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.485660Z digest=sha256:ac155aa01f9cd43ebd657664cde6b23e83506a7f7c4f20ccccdd4b546a1cc647

Observation 1215aaac-4870-451d-8d1f-5006b754c1fa · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.539919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.539919Z digest=sha256:f7657503051e8a2ec3085fb3811c5f5887f84dbeab090432d9f2415406f080b5

Observation 27bba49a-c6ab-4884-8771-3e1f26187e7e · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.545689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.545689Z digest=sha256:31a061e37af504c6d82a1bff17d1442e3ec0c5adedc3427ea05a9c2e438a3504

Observation 0e8b2492-d185-4735-80b9-9012088498c6 · outbound

This paper cites Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.551402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.551402Z digest=sha256:439f1f5de735474a662d62924193c0c6eca411688f248a8bed1d45e4e0bea3b9

Observation 27ee5e9c-59fb-450c-ae64-9de5435ad704 · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.556605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.556605Z digest=sha256:e7885231142b325def3c4ae7dada049187135ba8a40095ca3f4eacfaf876bccc

Observation 8038c5c9-bb5e-49fb-8d72-70ff9035f7cf · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.560788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.560788Z digest=sha256:7f70c2e843dbc4a8d830465aa7cf0ec7e53ea738bbb4a74495bb4e33a1aebe27

Observation da7a92d6-6bd0-4504-8c55-5d2b3b2cf846 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.631020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.565094Z digest=sha256:cd3768157289299d41d28e371c7a359de86ae04267be69a1524da35a81cb3cad

Observation f5069de9-453b-47c6-9d16-d3e7e0ac01d1 · outbound

This paper cites QuRating: Selecting High-Quality Data for Training Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection QuRating: Selecting High-Quality Data for Training Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.569987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.569987Z digest=sha256:661c18872c524ae0a9e665ee95045b8305fea5bb101223a7b27a2a4d40f8a885

Observation 1648229f-8a32-450f-8ed0-2c1d03f64a6c · outbound

This paper cites Organize the Web: Constructing Domains Enhances Pre-Training Data Curation.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Organize the Web: Constructing Domains Enhances Pre-Training Data Curation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.576037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.576037Z digest=sha256:23137cbe59b148b1b4919cbc11746c7bdef16badaa40d2389e6f0808bede8aef

Observation 7d1e7c72-9bd7-49a3-8403-c1dcfea521f5 · outbound

This paper cites Retrieval Head Mechanistically Explains Long-Context Factuality.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Retrieval Head Mechanistically Explains Long-Context Factuality

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.581532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.581532Z digest=sha256:13cb6c03a73be5308c407efa8d8ccc85dd4ab3eb5a49cb443a40294a11177d40

Observation c1106b25-412c-4878-b6d7-a6f0c0abe522 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36:69798–69818, 2023.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Doremi: Optimizing data mixtures speeds up language model pretraining.Advances in Neural Information Processing Systems, 36:69798–69818, 2023

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.585807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.585807Z digest=sha256:c85b0b202cf59094573a8e2bca805db411c1cc8cd4a10bf08c6a5f72c337fbb7

Observation 2105e29b-e248-4728-8d25-f58d390110a2 · outbound

This paper cites Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Data Mixing Laws: Optimizing Data Mixtures by Predicting Language Modeling Performance

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.590686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.590686Z digest=sha256:59d477be7977429c7acff15e45084d11ade3668c7c720760db33e02684d50156

Observation 9088e316-d5e9-46c9-b83e-f793a793c9f1 · outbound

This paper cites Mates: Model-aware data selection for efficient pretraining with data influence models.Advances in Neural Information Processing Systems, 37:108735–108759, 2024.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Mates: Model-aware data selection for efficient pretraining with data influence models.Advances in Neural Information Processing Systems, 37:108735–108759, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.608675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.596269Z digest=sha256:497c36d3cd6b7a5c53b30552ee2865110ae4374cd6144aa38d0d6198bc163e12

Observation 727e857c-c636-488d-a5c8-3c70e7a52234 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.600908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.600908Z digest=sha256:60edd75f16c6ac1ba7ef076afc193435db344b7abae6acc2573a1854f19926da

Observation 67b5ae53-835e-4bfe-bd07-633422466e4b · outbound

This paper cites DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection DecorateLM: Data Engineering through Corpus Rating, Tagging, and Editing with Language Models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T22:23:36.894507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.606615Z digest=sha256:70e6e72076e0f3417d15084ac16d9cbc94e6d6248a861b6408653fa44162cadb

Observation e409344f-83e1-4c8e-b326-ab841ba8882b · outbound

This paper cites Attention Heads of Large Language Models: A Survey.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Attention Heads of Large Language Models: A Survey

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.611340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.611340Z digest=sha256:a32fd7091d7131139bfdd87708785bcd9c6af152f11ae0155dbc3b093e59cd87

Observation 37fdae11-9ece-4c3e-95b9-a05642bae739 · outbound

This paper cites AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:36.616305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.616305Z digest=sha256:0e3067e58c542d17736df768c99f8b89316093dea03f8421ed9246fb4701ded5

Observation 8c57fe57-fa0d-4e2a-afdb-205bbbc40cc4 · outbound

This paper cites Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Focus Directions Make Your Language Models Pay More Attention to Relevant Contexts

Reference 55

Resolution
malformed identifier
no resolver link, observed 2026-08-15T22:23:36.648696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:36.648696Z digest=sha256:7891f1595ba3b3578df0f4d600f3f053f84c54fd3f42ae3f1dc29541463616e1

Observation d61d900d-6071-44fc-825a-f19e204938d7 · outbound

This paper cites "This is a test string.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection "This is a test string

Reference 57

Resolution
metadata mismatch
raw_fallback, observed 2026-08-15T22:23:36.832878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.688883Z digest=sha256:65fa9f9a7ef9b2b5d88f82b0573435a8b5268da5c4cae4b200ef37818549e58f

Observation c016396b-fa74-44e1-822b-b2eef5aba8ad · outbound

This paper cites 26 Figure 18 The cloud maps of the data selected by AttentionInfluence and FineWeb-Edu Classifier, respectively.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection 26 Figure 18 The cloud maps of the data selected by AttentionInfluence and FineWeb-Edu Classifier, respectively

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:23:37.593637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T22:23:36.717737Z digest=sha256:b068e05c055e21f0e1ab5f57452689afe679a35f0f8adde51f05994533b6a2c0

Observation 6fea8a0a-3d62-4e62-9dbb-15a270189090 · outbound

This paper cites an unresolved cited work.

AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T22:23:35.361541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:23:35.361541Z digest=sha256:5e1481b894786c602f4ede6bdabe5ffaf919f95d04c6a4d3330f32ca94fdf28d

Pith citing papers

Observation 2ea950b2-81b0-4595-bed0-f55952963bc9 · inbound

Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality cites this paper.

Hubs or Fringes: Pretraining Data Selection via Web Graph Centrality AttentionInfluence: Adopting Attention Head Influence for Weak-to-Strong Pretraining Data Selection

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.035631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T12:57:26.458109Z digest=sha256:3ba17ab41b989a717f9f178d58744a48cc840b167a96dcd7ac9249e98e486518