Pith. sign in

Paper Citation Record · LEDGER

TopK Language Models

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21468 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:39.220992Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 789d64b4-69c1-41fb-80dd-51772945f630 · outbound

This paper cites How can we be so dense? the benefits of using highly sparse representations, 2019.

TopK Language Models How can we be so dense? the benefits of using highly sparse representations, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.922624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:19.862785Z digest=sha256:a7849c5b29c75fd78e79d4405a96a948357fddfb720724381ed47189777270d0

Observation 252a90f8-e166-4ccf-8aa8-0427bd7175dd · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

TopK Language Models Generating long sequences with sparse transformers, 2019

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.787577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.787577Z digest=sha256:6d41960b22f14c0953f2fdcc9f2bf56094089c58c87f341cd48cecfa6f1f50ea

Observation 674b7941-d134-4ec2-b182-617c4b245cb5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TopK Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.894657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.894657Z digest=sha256:08752f98c3a2220cdfc5601b17e27d35655ee82477bb520b9bc3554419b68fcd

Observation 7a6bcd59-399b-4e81-a450-405ea058ef91 · outbound

This paper cites [Full Post] Progress Update #1 from the GDM Mech Interp Team.

TopK Language Models [Full Post] Progress Update #1 from the GDM Mech Interp Team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.790819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076386Z digest=sha256:5f3ae7444c4d6e458c9dd34ebfb41d5049adb0f3d02a25026dc431a03af2e82d

Observation ba9b874e-ce44-4514-9642-ba67d7764f98 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models, 2023.

TopK Language Models Sparse autoencoders find highly interpretable features in language models, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.158355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.158355Z digest=sha256:79648a9b04d368a7d76f2d72e5814e7a7ecfaf625ec081a159d3e1837551a2bb

Observation f0204e56-9a10-4e7c-947b-1d805f936fb0 · outbound

This paper cites Rigging the lottery: Making all tickets winners.

TopK Language Models Rigging the lottery: Making all tickets winners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.685373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.321207Z digest=sha256:df7b1951f17c89a06b483d0c703658b1522ca61257cd67ad22ee70b8eae5047c

Observation 30ea9f6d-00f7-4f49-964f-0f3dc6a16879 · outbound

This paper cites Sparsity in transformers: A systematic literature review.

TopK Language Models Sparsity in transformers: A systematic literature review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.560982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434488Z digest=sha256:b4707f9f817591d6b23b3e253a9f337dda9dd9d55da51ebc5cc0bfe6c4a68561

Observation 72d4a652-94cd-4b34-a2aa-d89f288ff9b0 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

TopK Language Models Detecting hallucinations in large language models using semantic entropy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530869Z digest=sha256:2963621f0042cd5693eaed393ec82e17eca31fd6c77158fe40f6ea4b7879cae3

Observation 6f137638-06ec-406b-9ce0-061a3bf57b53 · outbound

This paper cites Scaling and evaluating sparse autoencoders, 2024.

TopK Language Models Scaling and evaluating sparse autoencoders, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.452327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.680217Z digest=sha256:2fed998841f92bb650dd813739195ea3827796142bc18225724d848880c50847

Observation f4e0c28a-7209-4240-955b-333c6e0ea7da · outbound

This paper cites The language model evaluation harness, 2024.

TopK Language Models The language model evaluation harness, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.370392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.732571Z digest=sha256:41ca3e99a1689308fcaf0f2216fcdc7bbe294f0c30d51f5f5e2708bca409c9aa

Observation ee4dd635-db57-410b-9324-1787e68f3a5c · outbound

This paper cites Causal Abstractions of Neural Networks.

TopK Language Models Causal Abstractions of Neural Networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.255996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.794053Z digest=sha256:c3678f2fef32b98737c6b469876515757083836f1b2ae5a5dfe8fd507003aab0

Observation 04f5fb92-8a51-49c8-b8d4-d99b82567568 · outbound

This paper cites The State of Sparse Training in Deep Reinforcement Learning.

TopK Language Models The State of Sparse Training in Deep Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.158063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:35.872432Z digest=sha256:08aa7e61e2eed54d5d10a561b7359dd0d15f7b5886185b175c76ef3c15a12f8a

Observation aa35ef23-a1e2-4365-8576-b6daf9a5c923 · outbound

This paper cites Memory-efficient transformers via top-k attention, 2021.

TopK Language Models Memory-efficient transformers via top-k attention, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.047952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032941Z digest=sha256:edf5fb47d100dfc6e648c368d1f33929cc558f39b71969f2d132939a67516453

Observation 3f7984cd-049b-4bda-86f3-9384fd31b028 · outbound

This paper cites Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024.

TopK Language Models Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.933604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.096247Z digest=sha256:2515a983ace18904e23a7de69d6a35e0ab70f2dab4deb688eac7e21e71f2b215

Observation cdb4a400-794b-422a-b9db-23e9e5554868 · outbound

This paper cites Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba.

TopK Language Models Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.826679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.144498Z digest=sha256:497732b693ee92f283b1a092e15da1711f80d1558497d87ed2c80f37865881f6

Observation 93a23817-5ae8-4063-9a61-cc7a57c7407d · outbound

This paper cites Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks.

TopK Language Models Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:39.565510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.250242Z digest=sha256:b0d4dd89a8f909d67da1ef3386faba754588b7eab235e9ab23d454031bd71999

Observation 431d5c70-f9cd-434d-b265-d0c8d173ceb6 · outbound

This paper cites How llms learn: Tracing internal representations with sparse autoencoders, 2025.

TopK Language Models How llms learn: Tracing internal representations with sparse autoencoders, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.727068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.341495Z digest=sha256:55ce1f288d72c0d7ea662f3a1a435c3c31a4644879295b5e6b272140d8e0b163

Observation d3c0bed0-dcef-4bad-a2a3-7229280fc0ad · outbound

This paper cites Sparse is enough in scaling transformers, 2021.

TopK Language Models Sparse is enough in scaling transformers, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.615792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403525Z digest=sha256:5b9a345a619339ab47aa24315586005cee636473511dcdd4467048d01dfbf44b

Observation d7578cb5-abb8-45ff-b7d9-85726b9abc70 · outbound

This paper cites Top-KAST: Top-K Always Sparse Training.

TopK Language Models Top-KAST: Top-K Always Sparse Training

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:31:39.411734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509531Z digest=sha256:d6063210741f47c2f4c380124f4efebe84288b0c7ec07e5e32fb9c36be3ecd20

Observation 4f91072b-f53e-4f6f-b9bf-d3d8d7b64d13 · outbound

This paper cites Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025.

TopK Language Models Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.505754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629380Z digest=sha256:10312b3bd998761c5cd786bd5486d975cfae7816530444f3c4bdc6ffc30ce528

Observation 1d8d7c24-9783-465d-9c80-be28083f430a · outbound

This paper cites Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025.

TopK Language Models Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.327447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737124Z digest=sha256:26c0e62bd8ba149f20b3ec2e4778de289ab759753ac11281d6cd60dd4de1c5ae

Observation 88cbb3b3-8ac3-4ba0-8158-3b1251dba858 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

TopK Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842439Z digest=sha256:30b19ec636aa94ecdcad7d209332f366c9608a8399407689777ea67a60436246

Observation 3edfc9c7-94f6-4281-a590-880224887ad8 · outbound

This paper cites Soft Threshold Weight Reparameterization for Learnable Sparsity.

TopK Language Models Soft Threshold Weight Reparameterization for Learnable Sparsity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.002177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925606Z digest=sha256:bf4cbd129c6437f7e46b65168d393204406d1d77632cc40f49c6475749bc73ee

Observation a05d54fd-ae7b-4413-b870-4cecc3413d54 · outbound

This paper cites Sparse autoencoders do not find canonical units of analysis, 2025.

TopK Language Models Sparse autoencoders do not find canonical units of analysis, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.807089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983433Z digest=sha256:3d02bc33a313588b52c6b3c7ed1f0c428db3acc5a499229320481e4a0874b621

Observation 5f4fba51-bb4a-4ebd-9a82-f17b303fec40 · outbound

This paper cites Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024.

TopK Language Models Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.056013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.056013Z digest=sha256:f56fd76e3b20e7eea260162dc284526a2deac6561fb218f6128ddf5f30003b20

Observation 0e04cee2-46f0-46d1-ade5-282892edb508 · outbound

This paper cites Decoupled Weight Decay Regularization.

TopK Language Models Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178870Z digest=sha256:67385f18a5c8bec041b712323f59784633630a7adf98a3ddc93170222140559e

Observation dfa20b76-4411-44f4-807f-6be32170786e · outbound

This paper cites Learning sparse neural networks through l_0 regularization.

TopK Language Models Learning sparse neural networks through l_0 regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.623753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243550Z digest=sha256:b9bbf9d1e9fbe853c4fe22ff7c54e507abb474572a0ba30843fc102df6063b2f

Observation 8a792f02-b2c6-45e8-b471-bb1f4f042e41 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

TopK Language Models Fineweb-edu: the finest collection of educational content, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336793Z digest=sha256:0089a34028976afeb64addb68de8cf71ae0e9699ff1eda6df8eb11d8f99a9bce

Observation c81c18f1-a5e1-419b-878a-9a4b2a29115b · outbound

This paper cites Winner-take-all autoencoders.

TopK Language Models Winner-take-all autoencoders

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.399846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.409944Z digest=sha256:ca4a0ad6c34a481191c1b353040e89f6e4b326e4cbe0cf80f74a09befa0856a1

Observation cf1a13db-df4e-4f85-8d87-3a1f51a45d56 · outbound

This paper cites Pointer Sentinel Mixture Models.

TopK Language Models Pointer Sentinel Mixture Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.252639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516145Z digest=sha256:88cc1fb902bf249806b5a60717156a441d6be0d0dabeca62dbfd692f607fe895

Observation 6cc3d4c0-c29f-45dc-b26e-406dba71cc93 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

TopK Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.067358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572894Z digest=sha256:d9946f75d70e1f61f15b7968bd4bae72ef3b4119ff8ecafbc783fd230a47e0d7

Observation 69e9c625-3b6b-45ee-883a-cc68e008da50 · outbound

This paper cites Nguyen, Madeleine Gibescu, Antonio Liotta, and et al.

TopK Language Models Nguyen, Madeleine Gibescu, Antonio Liotta, and et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.886349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620117Z digest=sha256:966552b3e7c4088e46f8f40040490c96cfb0ec3b49550a00f04989c484f70328

Observation 2a719d09-833e-4914-8ee3-0b74cd9203b5 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TopK Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.688817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.666381Z digest=sha256:37cadd106c7341de8ca707092e944a063ea337f81271ac3f67f0deef5a2f5440

Observation 6dd0f983-9ab7-4e9c-8faf-5f2e127b6e35 · outbound

This paper cites Sparse autoencoders trained on the same data learn different features, 2025.

TopK Language Models Sparse autoencoders trained on the same data learn different features, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.469546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.703915Z digest=sha256:6d049dd9daa7254ff72acdd13f448cb30a53f6891f147eb4b9ee7a2738a84ba4

Observation 0754bfa7-a587-49c0-ab97-1f49c4ce81b1 · outbound

This paper cites Automatically interpreting millions of features in large language models, 2024.

TopK Language Models Automatically interpreting millions of features in large language models, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.268686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798385Z digest=sha256:953909458be80e96e837667071452d28547e5efd9fbabddcb10ab485c50eaab5

Observation a1125760-f628-4ca2-9ebb-4d6f608633bf · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

TopK Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.035447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:37.854065Z digest=sha256:21f5b6bcd344f729c2498d540b6c419b72374af79ae8e86adadd14853443ef00

Observation d02c0341-8289-4f6a-9e93-7c948279a759 · outbound

This paper cites Improving dictionary learning with gated sparse autoencoders, 2024.

TopK Language Models Improving dictionary learning with gated sparse autoencoders, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.972247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.972247Z digest=sha256:f3beea778390ff6b6dda6685a4faf12b161104ce3bc62f15fd90597da824b433

Observation acedd645-2f42-434a-8ef1-aebff0af90bb · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

TopK Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.142907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.142907Z digest=sha256:6a2d79683718ebad5cfa9ad7d2726611454c66438c8091ee30d01de1f4f94e88

Observation 8f6dab73-abb4-4500-b8d5-f326759064a8 · outbound

This paper cites Taking features out of superposition with sparse autoencoders.

TopK Language Models Taking features out of superposition with sparse autoencoders

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.851527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258973Z digest=sha256:adea7930a2c5cf55fbfda4d5a6f92e52720116fd84b3036df7043a539cc5b035

Observation ed8b5f35-897e-4c8d-9f38-e00bfacbbac6 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017.

TopK Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370614Z digest=sha256:bf13379607a7ae5c6c3d61e1c4b4338b8e2271fad49bdfffb22b9211afae5485

Observation 1ddad81e-ea7a-41e9-b420-c9e5d64197bd · outbound

This paper cites A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025.

TopK Language Models A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.670077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.509917Z digest=sha256:ca55da67d14682abc6b00aef75b55b5e9bdbb77d82d910a4a99cec4061a2d7d3

Observation 9c0496bf-71fc-42ea-a76c-b215fc6f7781 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

TopK Language Models Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.680344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.680344Z digest=sha256:1efa5c7e556e8310b40363961019f9601820cef23baca98c102dd74042ab318d

Observation e1b8015f-8dca-48ee-a6d5-94e0197c7317 · outbound

This paper cites Daniel Freeman, Theodore R.

TopK Language Models Daniel Freeman, Theodore R

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.462326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.794923Z digest=sha256:26405a9f73eec10f50d96d943be18424ef55161e92b0b5d8b6f3b3edab995b59

Observation 4e902680-9768-47c0-aae7-a1a9b5f98c7c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TopK Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.890079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.890079Z digest=sha256:0c383dbaae2f33765f0f74b788a3035e23ff18f02c117186704f1e63be9a7657

Observation 3ea79e63-1c97-4e4a-ad6c-80dc996c16ab · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

TopK Language Models Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.282755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:38.998159Z digest=sha256:f5f3302c5797ae35d7caa83ef2b91bc60f5b8a87d8bae5aad8ad80ea86096344

Observation 373449a0-f1a3-45e0-b63a-24c6155ab499 · outbound

This paper cites Tracking the feature dynamics in llm training: A mechanistic study, 2025.

TopK Language Models Tracking the feature dynamics in llm training: A mechanistic study, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.093824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.060459Z digest=sha256:09bbb0b586d4c7f12c310e61879605b91892a911c033f8a0f21deeba09e16bfe

Observation b26d6674-e1c9-49fd-a3d2-63892221c21a · outbound

This paper cites STEP: Staged parameter-efficient pre-training for large language models.

TopK Language Models STEP: Staged parameter-efficient pre-training for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.921354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121952Z digest=sha256:36e6b626873c2f3030d74e68b35ab2b14384ae141b897cb9cdbeb806f948f2d1

Observation 4a13895a-5f6b-42e3-a2c2-6056c77dd21a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

TopK Language Models HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.727732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T22:31:39.176817Z digest=sha256:cf94127a6b0b7f942763270136ff4a1d56761b0592c15380c1986af1c37150d4

Observation 757b2521-884d-4004-a072-b40518b8c712 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TopK Language Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:31:39.220992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.220992Z digest=sha256:0f83bbd6ce7caf6337edabc37b1d1efc0517e8055dff309462611e98e266c99c

Pith citing papers

No inbound Pith citation observations are available.