Pith. sign in

Paper Citation Record · LEDGER

On Linear Representations and Pretraining Data Frequency in Language Models

As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 4 inbound Pith citation observations for arXiv:2504.12459.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12459 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:37:36.423808Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T04:52:44.489356Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:06:17.562742Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact3
  • verified fuzzy15
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9be8aa00-ec8e-4c84-a732-a1276b6a6208 · outbound

This paper cites To code or not to code? exploring impact of code in pre-training.

On Linear Representations and Pretraining Data Frequency in Language Models To code or not to code? exploring impact of code in pre-training

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.117127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.219091Z digest=sha256:ca0eac95cbf39a09c8fb35cc023b5b5388acf4cf5ce2e3cc0d91da42d3c93a28

Observation 02a38762-7b87-431e-b172-30f41c838d6e · outbound

This paper cites Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers.

On Linear Representations and Pretraining Data Frequency in Language Models Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.222599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.222599Z digest=sha256:aac67bb9b5ef61b829f1363add942eabd94574fdf66fdc8b5eca9f59d352259f

Observation 724aa831-5196-4c4f-aeb0-9e158903b4bb · outbound

This paper cites Interpreting Neural Networks through the Polytope Lens.

On Linear Representations and Pretraining Data Frequency in Language Models Interpreting Neural Networks through the Polytope Lens

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.225333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.225333Z digest=sha256:4c93a4eb4efd82dcbdb0bfd8b88a4e85b8b5bd518b6992f6769b5b9a244977d2

Observation 65ade467-bb03-444d-a049-8e60da7558cc · outbound

This paper cites Membership inference attacks from first principles.

On Linear Representations and Pretraining Data Frequency in Language Models Membership inference attacks from first principles

Reference 4

Resolution
verified exact
raw_fallback, observed 2026-08-16T12:37:36.734597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.229182Z digest=sha256:bf58dabea066a819b823c4a2ee3ab87d144b49b77b7ac1b2209f5e9f3430d37e

Observation 690dd01c-c5c6-44ff-bf86-1a0bb0f9daa4 · outbound

This paper cites Quantifying memorization across neural language models.

On Linear Representations and Pretraining Data Frequency in Language Models Quantifying memorization across neural language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.232547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.232547Z digest=sha256:37348a620b1d507072846531510d6c8f5dc704e065fb4bd811d2d97f3a4d252d

Observation abe70b6b-8a87-4ac0-9423-9ec506a8e51b · outbound

This paper cites How do large language models acquire factual knowledge during pretraining? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024.

On Linear Representations and Pretraining Data Frequency in Language Models How do large language models acquire factual knowledge during pretraining? In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.098171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.235649Z digest=sha256:25c7c17d36c48744f4573901f781bfcd1e7709bb1dd7ada17028e28f610dea50

Observation 82581bb6-349b-49d6-9b5a-b1a531e61b08 · outbound

This paper cites Identifying Linear Relational Concepts in Large Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models Identifying Linear Relational Concepts in Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.239229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.239229Z digest=sha256:31ecad40c8e5260150a4f271c266494ec11ff4daa362c16e6d30fbe38d49788f

Observation 5d208efe-cb70-4887-a0ca-c1b5a09a3ad8 · outbound

This paper cites Recurrent neural networks learn to store and generate sequences using non-linear representations.

On Linear Representations and Pretraining Data Frequency in Language Models Recurrent neural networks learn to store and generate sequences using non-linear representations

Reference 8

Resolution
verified exact
doi, observed 2026-08-16T12:37:36.572317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.242809Z digest=sha256:fad09c53ce0522a7e589494c598ad6d858760abb6e67b564b2ad16bba071128b

Observation 296226ad-19b5-4b10-944c-4f34e5d2946c · outbound

This paper cites Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals.

On Linear Representations and Pretraining Data Frequency in Language Models Amnesic Probing: Behavioral Explanation with Amnesic Counterfactuals

Reference 9

Resolution
malformed identifier
no resolver link, observed 2026-08-16T12:37:36.245789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.245789Z digest=sha256:92f0ebeda87b99bff3e377583f70b95cb89776b3424b3b3209ebe3ce97ade997

Observation c29e93ab-2b3d-42e3-9d62-b4b8eb55e9a7 · outbound

This paper cites Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions.

On Linear Representations and Pretraining Data Frequency in Language Models Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.248759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.248759Z digest=sha256:80340fc7b7ee06e581fc66e31ac6b6bc8729c3dc5f5afbb9a2c809ebc2d8d3d5

Observation dff16e97-9766-4084-9ddf-47eae3b4dde7 · outbound

This paper cites Smith, and Jesse Dodge.

On Linear Representations and Pretraining Data Frequency in Language Models Smith, and Jesse Dodge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.251685Z digest=sha256:46e3b178422de18693c7360975d2a2187ce80b24af5c7ded7f5206421a398426

Observation 2128fed1-42a8-4911-864c-38476d366370 · outbound

This paper cites A mathematical framework for transformer circuits.

On Linear Representations and Pretraining Data Frequency in Language Models A mathematical framework for transformer circuits

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.255407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.255407Z digest=sha256:2530bb6ec97b4cc9bba4d8ae50a6f9f5bf9762e7d03aa1ac14195895f1f091be

Observation f9573e36-13d2-474b-9126-440c4b4dfe71 · outbound

This paper cites T - RE x: A large scale alignment of natural language with knowledge base triples.

On Linear Representations and Pretraining Data Frequency in Language Models T - RE x: A large scale alignment of natural language with knowledge base triples

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.067881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.258400Z digest=sha256:e8fa87401dd481d1f1240606661de82dd54630922dae0d9d3621fbec9e6c0dc5

Observation b27cfc5b-8070-4529-8f6b-b4934ec0c0ac · outbound

This paper cites Towards Understanding Linear Word Analogies.

On Linear Representations and Pretraining Data Frequency in Language Models Towards Understanding Linear Word Analogies

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.261244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.261244Z digest=sha256:86ccbf8ef3eb8ff19765316afafeee8a01854eb9eed042145cf195d5dd324be8

Observation ec363ae7-6085-4fd0-8ed9-1d9c0af324bc · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

On Linear Representations and Pretraining Data Frequency in Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.264651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.264651Z digest=sha256:05b7f0ed8dfc2ee0c31693d94c010c822c0e2fc832bf9214f4d4d81cd239e778

Observation 16ecfce0-7101-4a7e-b91e-ef1b3bd77a7f · outbound

This paper cites Scaling and evaluating sparse autoencoders.

On Linear Representations and Pretraining Data Frequency in Language Models Scaling and evaluating sparse autoencoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.268775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.268775Z digest=sha256:2df71314a9fd0214fd014d02cafeae2fafb764ad2f2e1b47e0ad5b6b5c990d3e

Observation dacfa710-47cd-4a08-a7fc-71c7dafc86e4 · outbound

This paper cites What can transformers learn in-context? a case study of simple function classes.

On Linear Representations and Pretraining Data Frequency in Language Models What can transformers learn in-context? a case study of simple function classes

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.048571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.272216Z digest=sha256:2406e88b9677931532117ac9feb82838520bdb5b5fd7c0b55e84ef4f8868c9c6

Observation 8011d74e-f40c-4da3-a142-ecbac4602b57 · outbound

This paper cites Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn`t.

On Linear Representations and Pretraining Data Frequency in Language Models Analogy-based detection of morphological and semantic relations with word embeddings: what works and what doesn`t

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.276369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.276369Z digest=sha256:5398d5fdd2107edf7bd56d673449144abcb9277a653d247ba06c1628e313a456

Observation 0f6d728a-31b0-4b19-b183-0614e70e8c81 · outbound

This paper cites OLM o: Accelerating the science of language models.

On Linear Representations and Pretraining Data Frequency in Language Models OLM o: Accelerating the science of language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.280450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.280450Z digest=sha256:e040d1a576afe9cabda6a801d2a52b1e8880f7cc2b6672ec6158c60d1446cb7d

Observation 61de6d86-90cd-4feb-8723-69a9970c8ba9 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-16T12:37:37.036827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.284194Z digest=sha256:12d51ec00063b05a845e9168a58034bd95b99e193ca9486181baade4f073f950

Observation fc1f87f2-39a9-495b-94b9-fdc8f5cd209b · outbound

This paper cites In- Context Learning Creates Task Vectors.

On Linear Representations and Pretraining Data Frequency in Language Models In- Context Learning Creates Task Vectors

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.287788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.287788Z digest=sha256:849d4e83b42b1cb5fe9ffa1fda02050cd8fe3179f2c118dfbc47e636195e5710

Observation 935a2fb3-0694-45eb-8236-cd6caac42e6d · outbound

This paper cites Linearity of relation decoding in transformer language models.

On Linear Representations and Pretraining Data Frequency in Language Models Linearity of relation decoding in transformer language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.292061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.292061Z digest=sha256:89f12d0eed0479a321c497e393d0dc69567d56b05bd6a4648351e3602c42f952

Observation ef2593e1-8db3-4dca-af3f-113b8d36f760 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models.

On Linear Representations and Pretraining Data Frequency in Language Models Sparse autoencoders find highly interpretable features in language models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.294958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.294958Z digest=sha256:a6aad7bf58eb3783f877cf39beac77b047f132538ef2fb0183e39d0bff08c2ed

Observation 1820c1dd-5bb4-4953-ad68-da5d79e0ca4a · outbound

This paper cites On the origins of linear representations in large language models.

On Linear Representations and Pretraining Data Frequency in Language Models On the origins of linear representations in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:37.010624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.297826Z digest=sha256:c18816933704a1784b9e34bb5cd53b11c3f6ec7894b817d5e9eb8bebf469606a

Observation 28864f5b-e906-4a6e-9f25-80469d09aa74 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 25

Resolution
verified exact
doi, observed 2026-08-16T12:37:36.528778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.300566Z digest=sha256:cd30f886f377fe151eb7838ae432d8634915efc3455d357bd1cb5c7487d4e0ad

Observation 63c8a6f8-3f32-4c64-969c-15cb5e33744f · outbound

This paper cites Multilingual reliability and semantic structure of continuous word spaces.

On Linear Representations and Pretraining Data Frequency in Language Models Multilingual reliability and semantic structure of continuous word spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.996973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.303582Z digest=sha256:11354c45fafe06f7de87db90e1818d8a4f785f91c29709fc67da45e982222247

Observation 9646f0d9-ee4c-458e-8ced-9bf0aaa0b3ba · outbound

This paper cites A pretrainer`s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity.

On Linear Representations and Pretraining Data Frequency in Language Models A pretrainer`s guide to training data: Measuring the effects of data age, domain coverage, quality, & toxicity

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.307341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.307341Z digest=sha256:918aea260b2ba8dd04927da6f7cd78225cc234222025d59d4fe0ab2fd10e97ed

Observation 80ebb938-d5d8-4a12-b775-374d0ec56c21 · outbound

This paper cites At which training stage does code data help LLM s reasoning? In The Twelfth International Conference on Learning Representations, 2024.

On Linear Representations and Pretraining Data Frequency in Language Models At which training stage does code data help LLM s reasoning? In The Twelfth International Conference on Learning Representations, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.314919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.314919Z digest=sha256:4557cafe8d75118ea4637adf1168e9d6dca892a3011dd89f4a649cdd1ffeb9db

Observation e08f2e27-05eb-4b4c-b93d-b3ca2e19131e · outbound

This paper cites When not to trust language models: Investigating effectiveness of parametric and non-parametric memories.

On Linear Representations and Pretraining Data Frequency in Language Models When not to trust language models: Investigating effectiveness of parametric and non-parametric memories

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.317695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.317695Z digest=sha256:70659a5b971dff6367c160d2da911e4f42fad7fbb20ad0ab8dabbdc71ce30656

Observation 5148ca2b-7301-431a-9bb8-06e8e79eb9c3 · outbound

This paper cites Embers of autoregression show how large language models are shaped by the problem they are trained to solve.

On Linear Representations and Pretraining Data Frequency in Language Models Embers of autoregression show how large language models are shaped by the problem they are trained to solve

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.320670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.320670Z digest=sha256:a58c10ef9262f696d6384bdc6c0c69de870ea02fddb41e5876e5052796121f80

Observation 7276c5e4-7105-4cd2-a26b-dbe878ecf07d · outbound

This paper cites Language models implement simple W ord2 V ec-style vector arithmetic.

On Linear Representations and Pretraining Data Frequency in Language Models Language models implement simple W ord2 V ec-style vector arithmetic

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.323521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.323521Z digest=sha256:9a8506beeb8ef4e1351c774e69f6c549d6ea2aa10572131ee65941fd3a0b5aad

Observation 07fb214a-a2e4-4e3b-9bf2-a3ce395845e2 · outbound

This paper cites Efficient Estimation of Word Representations in Vector Space.

On Linear Representations and Pretraining Data Frequency in Language Models Efficient Estimation of Word Representations in Vector Space

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.326920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.326920Z digest=sha256:9233afc7b2b94c99d7ea26157562c88c240ff18219cc0386d3936d8b4996f00a

Observation d423fe4a-a934-4b57-b535-2dd7941ac934 · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

On Linear Representations and Pretraining Data Frequency in Language Models Distributed representations of words and phrases and their compositionality

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.970610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.330488Z digest=sha256:a1e1afb9bf8b11765b0185a526443b6ee43403663f35bd52d269bec029577360

Observation c8898bc0-9f6a-4d58-8c0b-a8504aaba89c · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.333747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.333747Z digest=sha256:8d467e6c7f98ae8ff2a9c99fe68f4841b92ec3539a339323ed5329fae51c6e6d

Observation 002554d9-8b63-4da5-9dec-f759a6602b8c · outbound

This paper cites Zoom in: An introduction to circuits.

On Linear Representations and Pretraining Data Frequency in Language Models Zoom in: An introduction to circuits

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.959877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.336933Z digest=sha256:e39d8eee6e57e42a1671bda7bea70e73146f112394768398effcf7da7ea53cb2

Observation f021bf14-4dde-4a74-9088-6eb57540e730 · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

On Linear Representations and Pretraining Data Frequency in Language Models Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.341261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.341261Z digest=sha256:831d15ce0f5f73eab8b06bfdb1a9902e2940136e5a85222374c5cf534642249e

Observation c529b3bc-2e8f-452b-bdf4-4a5c14891550 · outbound

This paper cites Learning Hierarchical Structures with Linear Relational Embedding.

On Linear Representations and Pretraining Data Frequency in Language Models Learning Hierarchical Structures with Linear Relational Embedding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.940678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.344829Z digest=sha256:515b644738c9a7fe0460ba28b53131e64f5a3912ae765b905f6f802321987b0b

Observation 059f7a5b-ed0e-4560-abea-b6611f4a5ae9 · outbound

This paper cites The Linear Representation Hypothesis and the Geometry of Large Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models The Linear Representation Hypothesis and the Geometry of Large Language Models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.929566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.349063Z digest=sha256:29df721e276cf2545ab377446d96e56b15f0204865091f2f808e461b3e8bb505

Observation 78bf24e7-b5fc-47f4-98f7-6dbc93b669ab · outbound

This paper cites G lo V e: Global vectors for word representation.

On Linear Representations and Pretraining Data Frequency in Language Models G lo V e: Global vectors for word representation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.352815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.352815Z digest=sha256:fa10be47c7904d4044cca837c3a1cf65112a176e31b74b7f2eaffc6810700c41

Observation e8e3a177-e391-4a5d-b57d-a290d40dad70 · outbound

This paper cites Null it out: Guarding protected attributes by iterative nullspace projection.

On Linear Representations and Pretraining Data Frequency in Language Models Null it out: Guarding protected attributes by iterative nullspace projection

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.917292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.356294Z digest=sha256:4963280e25e6ea113b272c1c909ddf5f0c2f8970b110bbe3bcee343c9e2495e5

Observation a78393f9-dd28-4252-89bc-4b33ca76a95d · outbound

This paper cites Impact of pretraining term frequencies on few-shot numerical reasoning.

On Linear Representations and Pretraining Data Frequency in Language Models Impact of pretraining term frequencies on few-shot numerical reasoning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.359724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.359724Z digest=sha256:7769062193433a4c1ed3bb1a4a6a446f888af51a3c4620c37d48b2c9988ac222

Observation 77363836-c1b2-4098-b3eb-386f16046404 · outbound

This paper cites Backtracking mathematical reasoning of language models to the pretraining data.

On Linear Representations and Pretraining Data Frequency in Language Models Backtracking mathematical reasoning of language models to the pretraining data

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.905434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.362674Z digest=sha256:b39ac1da9fbd35d213404a99c8bf76da49ebf32f78d0f230055b60022875f6a2

Observation 439fd0fe-b768-4db5-8271-0a4eea112db9 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

On Linear Representations and Pretraining Data Frequency in Language Models Steering llama 2 via contrastive activation addition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.365962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.365962Z digest=sha256:352b2828e48d477e80eb123a5fbfa4fe1eea434c06c4525ab4eea0d330ec4747

Observation f2046014-23a3-4e24-a48e-82f066f28c8c · outbound

This paper cites Salton, A.

On Linear Representations and Pretraining Data Frequency in Language Models Salton, A

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.369689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.369689Z digest=sha256:6dcf0d868dec8038868c9113dba08efc2f1353870da7261a6d709702d1edd95e

Observation dd67bddf-8ca1-4c1e-b87e-44f74ebb32a4 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.372202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.372202Z digest=sha256:d0f7429779f58b728ab3758a4d0cad08d38c0b5fe878881f1d7e9c6213483e04

Observation 3f630405-b32b-49dd-a1de-c818931f275e · outbound

This paper cites The bias amplification paradox in text-to-image generation.

On Linear Representations and Pretraining Data Frequency in Language Models The bias amplification paradox in text-to-image generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.375367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.375367Z digest=sha256:5723b42a2c6aa82147f38300abdb671c0f5f7c001d613adf9c80348746ae2932

Observation f37685d7-e8c3-4065-99d2-e447deaff3da · outbound

This paper cites Detecting pretraining data from large language models.

On Linear Representations and Pretraining Data Frequency in Language Models Detecting pretraining data from large language models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.378430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.378430Z digest=sha256:29808ced09f062cdbea8c6c109f30a9f04d64be667d3bb41fa4d64d3a66fbe1f

Observation 707bccd1-206e-49a4-ad5f-103f29e67cc9 · outbound

This paper cites Membership inference attacks against machine learning models.

On Linear Representations and Pretraining Data Frequency in Language Models Membership inference attacks against machine learning models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.381493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.381493Z digest=sha256:fe1726b9bb45d3d23c6462bda2cc3f49ff5c7f36775caba2d825a85004513a21

Observation 8bdb9137-e7d0-4ed3-b297-acb6ea1f6d2a · outbound

This paper cites The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models.

On Linear Representations and Pretraining Data Frequency in Language Models The curious case of hallucinatory (un)answerability: Finding truths in the hidden states of over-confident large language models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.385139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.385139Z digest=sha256:3de09923ce06cf4d6612aaac00f077ebaa238edd78c5c76648f78e97c20110a2

Observation 87c6a1e0-ac1d-4729-875f-966a3e12ef57 · outbound

This paper cites Dolma: an open corpus of three trillion tokens for language model pretraining research.

On Linear Representations and Pretraining Data Frequency in Language Models Dolma: an open corpus of three trillion tokens for language model pretraining research

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.388267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.388267Z digest=sha256:aecbfbdb8657da3e67a022e98c4282acaece688e4890a8308c265f6348343511

Observation c8182546-7004-4d5e-b677-b7a59e917a92 · outbound

This paper cites Extracting Latent Steering Vectors from Pretrained Language Models.

On Linear Representations and Pretraining Data Frequency in Language Models Extracting Latent Steering Vectors from Pretrained Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.390855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.390855Z digest=sha256:4ab9fe9c5ef0558922495b07ad475da9338108ce5910ddb78c94e407f3c14ef5

Observation 3bf08ac5-ee0c-47ea-af86-ba018a230d97 · outbound

This paper cites Formalizing and Estimating Distribution Inference Risks.

On Linear Representations and Pretraining Data Frequency in Language Models Formalizing and Estimating Distribution Inference Risks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.393954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.393954Z digest=sha256:09ff7d9b9f8b4eba26532303fafa34c5b1f71b247bcbde3b00c0abc249ac3dc0

Observation ab1dba8d-4911-47ac-b555-1fe3ce592346 · outbound

This paper cites Scaling Monosemanticity : Extracting Interpretable Features from Claude 3 Sonnet.

On Linear Representations and Pretraining Data Frequency in Language Models Scaling Monosemanticity : Extracting Interpretable Features from Claude 3 Sonnet

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.880815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.397117Z digest=sha256:26139117c692f740a65854f4dea73a21ca368f848514fdaeaa045d5291694b11

Observation c10e049a-c1a9-429b-a62f-f7d6d78e05fd · outbound

This paper cites Function vectors in large language models.

On Linear Representations and Pretraining Data Frequency in Language Models Function vectors in large language models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.399886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.399886Z digest=sha256:da8741cc232839b7687f0a54bea04cd2d3d2a76aaab062c1930884eb2b4ad1e9

Observation 980dfd8b-4b2f-4b44-866e-321f219ce919 · outbound

This paper cites GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model.

On Linear Representations and Pretraining Data Frequency in Language Models GPT-J-6B: A 6 Billion Parameter Autoregressive Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.402311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.402311Z digest=sha256:b3865365f529d22a73beeedeee47adc18f28c78f4378db2b201315a28cc0045f

Observation 2d105eac-e5ca-42f4-b4de-233b54968e97 · outbound

This paper cites Understanding reasoning ability of language models from the perspective of reasoning paths aggregation.

On Linear Representations and Pretraining Data Frequency in Language Models Understanding reasoning ability of language models from the perspective of reasoning paths aggregation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:37:36.855643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-16T12:37:36.405447Z digest=sha256:472992722ee6d3e2afb8ba3c5b0a0d35c043557ae2204a3456deddf5b2787ffe

Observation a2631398-a306-4511-9fc1-44dc66fb242d · outbound

This paper cites Generalization v.s.

On Linear Representations and Pretraining Data Frequency in Language Models Generalization v.s

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.408579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.408579Z digest=sha256:e76570e33cc0bbc1111c5785da2aad8b890035a693274994fe1a35476135c520

Observation 2f03f27a-ce68-4dbd-8a02-dd110def5698 · outbound

This paper cites Doremi: Optimizing data mixtures speeds up language model pretraining.

On Linear Representations and Pretraining Data Frequency in Language Models Doremi: Optimizing data mixtures speeds up language model pretraining

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.411183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.411183Z digest=sha256:33c9acecd36e866fd11516ae02efc6b470daad488d2d925e6de001bcfdf63012

Observation c1b8bedb-0c0b-4e4e-aca8-801ac0042f62 · outbound

This paper cites write newline.

On Linear Representations and Pretraining Data Frequency in Language Models write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.413796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.413796Z digest=sha256:a08bea2961bd5363cf4da3a4754c78f2df9c107a651abdf1d39d64e79b3a2cbf

Observation 10cc502b-2936-448a-a369-1dcf261fe555 · outbound

This paper cites @esa (Ref.

On Linear Representations and Pretraining Data Frequency in Language Models @esa (Ref

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.417232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.417232Z digest=sha256:ae2387d112aa6fdeb1b4a03a4b36e09087e52ab0ecfd8b2f546809f305c64f22

Observation 9f53c3e7-0937-4ab4-9c36-cbe7f4006e84 · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.420669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.420669Z digest=sha256:f1bbbe03c6fdd6659430bc1b72adec850a3a4816cbf93618eee59a995df5e309

Observation 4fd38eb4-93c0-4589-96fe-51cfee1bf51e · outbound

This paper cites an unresolved cited work.

On Linear Representations and Pretraining Data Frequency in Language Models Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-16T12:37:36.423808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:37:36.423808Z digest=sha256:3eacba6b2b783d9a08483a6a2cbb8bdf54bc7c1a5fff24930405076b14adcc36

Pith citing papers

Observation eb9371a3-71e5-430c-a26b-c3960bc65be3 · inbound

How Do Language Models Compose Functions? cites this paper.

How Do Language Models Compose Functions? On Linear Representations and Pretraining Data Frequency in Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:06:17.566205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T11:04:09.179536Z digest=sha256:e199e9ff34d1349bcbd4d6484a7f52208e278b12fe37be970dc8a7a74d37147d

Observation b4298fba-acae-4c14-886a-9dc5eb2345b5 · inbound

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space cites this paper.

Stories in Space: In-Context Learning Trajectories in Conceptual Belief Space On Linear Representations and Pretraining Data Frequency in Language Models

Reference 150

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:27:19.242265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T05:17:34.283917Z digest=sha256:9cb5979103c3f30867454734ada99139058d92dd91640263874766777baab9dc

Observation 72086d85-fd50-4d02-ae8f-6cd0aac35639 · inbound

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering cites this paper.

Graded Entity-Familiarity Readouts in Language Models: Polish Adaptation, Cross-Language Robustness, and Refusal Steering On Linear Representations and Pretraining Data Frequency in Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T04:52:44.489356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:52:44.489356Z digest=sha256:2c943517cc6cf4490f37c3c445d5c557766e10e865519f43686046309ff9be9c

Observation e530c710-5451-478f-b005-66aa643c563f · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining On Linear Representations and Pretraining Data Frequency in Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:05.090555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:05.090555Z digest=sha256:a81b4f1fad9746be7668c6b80b3db35c41d31c824c92921b93d87765c4547bd5