Pith. sign in

Paper Citation Record · LEDGER

TopK Language Models

As of 9 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 0 inbound Pith citation observations for arXiv:2506.21468.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21468 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:31:39.220992Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

49 of 49 outbound references displayed

  • verified exact1
  • verified fuzzy34
  • unresolved12
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 789d64b4-69c1-41fb-80dd-51772945f630 · outbound

This paper cites How can we be so dense? the benefits of using highly sparse representations, 2019.

TopK Language Models How can we be so dense? the benefits of using highly sparse representations, 2019

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.922624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:19.862785Z digest=sha256:252c4fa2aa2304cd1a28483c68e8a6f53c6c9aebdcb546504b0bf4e26e697267

Observation 252a90f8-e166-4ccf-8aa8-0427bd7175dd · outbound

This paper cites Generating long sequences with sparse transformers, 2019.

TopK Language Models Generating long sequences with sparse transformers, 2019

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.787577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.787577Z digest=sha256:b4a2391c2d30a0c9b6d369a0e14fcbe71597c7eff1e7ee7df740dd4f34550d85

Observation 674b7941-d134-4ec2-b182-617c4b245cb5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TopK Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:34.894657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:34.894657Z digest=sha256:26a3cc74f5f74415fc9c1f28fd974cce2cb52f7a3c2732c85dea6dcc0638c468

Observation 7a6bcd59-399b-4e81-a450-405ea058ef91 · outbound

This paper cites [Full Post] Progress Update #1 from the GDM Mech Interp Team.

TopK Language Models [Full Post] Progress Update #1 from the GDM Mech Interp Team

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.790819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.076386Z digest=sha256:7564db4386a0920cdbb5d428b219bff6338986a0405ceac79b370fa3a157a96c

Observation ba9b874e-ce44-4514-9642-ba67d7764f98 · outbound

This paper cites Sparse autoencoders find highly interpretable features in language models, 2023.

TopK Language Models Sparse autoencoders find highly interpretable features in language models, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.158355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.158355Z digest=sha256:d3ec71bea4761ace720ba1e36d0a7dfe582a541ce8215daf5b7f0d8c2b49be07

Observation f0204e56-9a10-4e7c-947b-1d805f936fb0 · outbound

This paper cites Rigging the lottery: Making all tickets winners.

TopK Language Models Rigging the lottery: Making all tickets winners

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.685373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.321207Z digest=sha256:af8c7365200ea88db0dce5481ccd1f461b4e3500c721c1be436d3324a55f332e

Observation 30ea9f6d-00f7-4f49-964f-0f3dc6a16879 · outbound

This paper cites Sparsity in transformers: A systematic literature review.

TopK Language Models Sparsity in transformers: A systematic literature review

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.560982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.434488Z digest=sha256:ca0f3044a1cc139773e5e86bf47ed455e50a3c162669a5e299dda9626414783f

Observation 72d4a652-94cd-4b34-a2aa-d89f288ff9b0 · outbound

This paper cites Detecting hallucinations in large language models using semantic entropy.

TopK Language Models Detecting hallucinations in large language models using semantic entropy

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:35.530869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:35.530869Z digest=sha256:c9348b0c76eeeefd206263e88182a5f58d8a2ebee154f146aec2a0548ae409b1

Observation 6f137638-06ec-406b-9ce0-061a3bf57b53 · outbound

This paper cites Scaling and evaluating sparse autoencoders, 2024.

TopK Language Models Scaling and evaluating sparse autoencoders, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.452327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.680217Z digest=sha256:1295be9b65af4f491fe6e4ffea282d9c4794e75da732ed6aa0b46682eb756982

Observation f4e0c28a-7209-4240-955b-333c6e0ea7da · outbound

This paper cites The language model evaluation harness, 2024.

TopK Language Models The language model evaluation harness, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.370392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.732571Z digest=sha256:2333610f34b9d24edfbb08d9afde2ce1e90da9812d53f424a3d8ab2037da62af

Observation ee4dd635-db57-410b-9324-1787e68f3a5c · outbound

This paper cites Causal Abstractions of Neural Networks.

TopK Language Models Causal Abstractions of Neural Networks

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.255996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.794053Z digest=sha256:a86c63e7e65496393a3a80057b288c222306fcf84a99d61433723a1778c62f4e

Observation 04f5fb92-8a51-49c8-b8d4-d99b82567568 · outbound

This paper cites The State of Sparse Training in Deep Reinforcement Learning.

TopK Language Models The State of Sparse Training in Deep Reinforcement Learning

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.158063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:35.872432Z digest=sha256:c7f9aaebb3b7b8911ef2b99ee912eb3f4826bb4d9ba9e0affff2e709b60bf796

Observation aa35ef23-a1e2-4365-8576-b6daf9a5c923 · outbound

This paper cites Memory-efficient transformers via top-k attention, 2021.

TopK Language Models Memory-efficient transformers via top-k attention, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:44.047952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.032941Z digest=sha256:4d7c87ce5754a09626a1f3f10426b9aa8c585a7430aba18444e5ac42ec671927

Observation 3f7984cd-049b-4bda-86f3-9384fd31b028 · outbound

This paper cites Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024.

TopK Language Models Llama scope: Extracting millions of features from llama-3.1-8b with sparse autoencoders, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.933604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.096247Z digest=sha256:55e1d3e9909f3fde3ef41c7f23dc6c1226a335e0f257fb255e73681593a4ef5c

Observation cdb4a400-794b-422a-b9db-23e9e5554868 · outbound

This paper cites Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba.

TopK Language Models Hindupur, Ekdeep Singh Lubana, Thomas Fel, and Demba Ba

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.826679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.144498Z digest=sha256:645bf28e272e6cb05107aa861836feefe83f1910abd49f1906a4e6246e4a1a72

Observation 93a23817-5ae8-4063-9a61-cc7a57c7407d · outbound

This paper cites Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks.

TopK Language Models Two Sparsities Are Better Than One: Unlocking the Performance Benefits of Sparse-Sparse Networks

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:31:39.565510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.250242Z digest=sha256:2adf2b16a7ef07a183c7c271f2490267797ae8b21313f6eadd5f08f879176888

Observation 431d5c70-f9cd-434d-b265-d0c8d173ceb6 · outbound

This paper cites How llms learn: Tracing internal representations with sparse autoencoders, 2025.

TopK Language Models How llms learn: Tracing internal representations with sparse autoencoders, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.727068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.341495Z digest=sha256:ed842cfc4e46ec87bdac1fe62a08e80b1cdfe620bb4945b7c6cf76164c941297

Observation d3c0bed0-dcef-4bad-a2a3-7229280fc0ad · outbound

This paper cites Sparse is enough in scaling transformers, 2021.

TopK Language Models Sparse is enough in scaling transformers, 2021

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.615792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.403525Z digest=sha256:977a91ae1a0f44091043aeb3543ca607c79e48b9b441b1d16cd946b825bb7c41

Observation d7578cb5-abb8-45ff-b7d9-85726b9abc70 · outbound

This paper cites Top-KAST: Top-K Always Sparse Training.

TopK Language Models Top-KAST: Top-K Always Sparse Training

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T22:31:39.411734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.509531Z digest=sha256:db695649839370d7bfba61840e7d02ccd7354c19b6446b28e420fefe3381d86c

Observation 4f91072b-f53e-4f6f-b9bf-d3d8d7b64d13 · outbound

This paper cites Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025.

TopK Language Models Saebench: A comprehensive benchmark for sparse autoen- coders in language model interpretability, 2025

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.505754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.629380Z digest=sha256:ee3438b0f5abee59f1186d5a0dea2b7a9aa9c13b142c18bf86d14f6130ed7288

Observation 1d8d7c24-9783-465d-9c80-be28083f430a · outbound

This paper cites Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025.

TopK Language Models Concept steerers: Leveraging k-sparse autoencoders for controllable generations, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.327447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.737124Z digest=sha256:cfb4c2d6579355e3659889a0c1048d5077aa8b751d67d6e7cc781388a24534d4

Observation 88cbb3b3-8ac3-4ba0-8158-3b1251dba858 · outbound

This paper cites Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation.

TopK Language Models Semantic Uncertainty: Linguistic Invariances for Uncertainty Estimation in Natural Language Generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.166195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.842439Z digest=sha256:e85e20d3047eed3d9567ee5dd46f62434d211533735e301dc91d402484c55a0e

Observation 3edfc9c7-94f6-4281-a590-880224887ad8 · outbound

This paper cites Soft Threshold Weight Reparameterization for Learnable Sparsity.

TopK Language Models Soft Threshold Weight Reparameterization for Learnable Sparsity

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:43.002177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.925606Z digest=sha256:b2654f55db35b11253e33a6b61bebce90e3ec43f13c5bd821b6774946eb632c1

Observation a05d54fd-ae7b-4413-b870-4cecc3413d54 · outbound

This paper cites Sparse autoencoders do not find canonical units of analysis, 2025.

TopK Language Models Sparse autoencoders do not find canonical units of analysis, 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.807089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:36.983433Z digest=sha256:50d3c6ba4caf8b95bdf1b3d76b4f728742573fbde6c3b7f3b790c76e495b9d42

Observation 5f4fba51-bb4a-4ebd-9a82-f17b303fec40 · outbound

This paper cites Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024.

TopK Language Models Gemma scope: Open sparse autoencoders everywhere all at once on gemma 2, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.056013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.056013Z digest=sha256:679ae33d18b55dfca61ea4e5ccad42dd055a38fabf8f17bdca80818fc3ad6e86

Observation 0e04cee2-46f0-46d1-ade5-282892edb508 · outbound

This paper cites Decoupled Weight Decay Regularization.

TopK Language Models Decoupled Weight Decay Regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.178870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.178870Z digest=sha256:1ee24beacd424609cd811ee695c93b9a0f744d677189eb999d0a1182888829c8

Observation dfa20b76-4411-44f4-807f-6be32170786e · outbound

This paper cites Learning sparse neural networks through l_0 regularization.

TopK Language Models Learning sparse neural networks through l_0 regularization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.623753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.243550Z digest=sha256:66551a72eefa3440e78dbc9b2fde3884edb64f661e254dc07f1303d4f37f6236

Observation 8a792f02-b2c6-45e8-b471-bb1f4f042e41 · outbound

This paper cites Fineweb-edu: the finest collection of educational content, 2024.

TopK Language Models Fineweb-edu: the finest collection of educational content, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.336793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.336793Z digest=sha256:40e9a8930e71364fecba091bd066f0178a1270212edc47c3d1d24cb7256164dc

Observation c81c18f1-a5e1-419b-878a-9a4b2a29115b · outbound

This paper cites Winner-take-all autoencoders.

TopK Language Models Winner-take-all autoencoders

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.399846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.409944Z digest=sha256:703c2b842ad3203ab974f74e2e72d37fddcbbd8a231684a0c11b7c8c2bac98ba

Observation cf1a13db-df4e-4f85-8d87-3a1f51a45d56 · outbound

This paper cites Pointer Sentinel Mixture Models.

TopK Language Models Pointer Sentinel Mixture Models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.252639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.516145Z digest=sha256:ec6181d81c8c84cba171425827a67a589eaa402430f0643aa516cf3dc068925d

Observation 6cc3d4c0-c29f-45dc-b26e-406dba71cc93 · outbound

This paper cites Can a suit of armor conduct electricity? a new dataset for open book question answering.

TopK Language Models Can a suit of armor conduct electricity? a new dataset for open book question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:42.067358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.572894Z digest=sha256:d3876bf8d1ea5c6944d08ea5cef7711cd8ead3d7dbfb205d04556c6cc28e5d22

Observation 69e9c625-3b6b-45ee-883a-cc68e008da50 · outbound

This paper cites Nguyen, Madeleine Gibescu, Antonio Liotta, and et al.

TopK Language Models Nguyen, Madeleine Gibescu, Antonio Liotta, and et al

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.886349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.620117Z digest=sha256:781d7018fdfeb9aac80b6281f682d894a9f9b247035b637c7ba88dfd56c4384a

Observation 2a719d09-833e-4914-8ee3-0b74cd9203b5 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TopK Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.688817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.666381Z digest=sha256:b992df4669e7fd2bf261e99f9187128bd37b1d0bd575ef60125b19dd4ea863d2

Observation 6dd0f983-9ab7-4e9c-8faf-5f2e127b6e35 · outbound

This paper cites Sparse autoencoders trained on the same data learn different features, 2025.

TopK Language Models Sparse autoencoders trained on the same data learn different features, 2025

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.469546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.703915Z digest=sha256:c075011872402be3bbb451f26d9bb9538bb47de562b14766b2d8759bbfb22d46

Observation 0754bfa7-a587-49c0-ab97-1f49c4ce81b1 · outbound

This paper cites Automatically interpreting millions of features in large language models, 2024.

TopK Language Models Automatically interpreting millions of features in large language models, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.268686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.798385Z digest=sha256:e2752001e0c679d2662c41aad970440874901e62fa9b0f0141e25c8b41e95baf

Observation a1125760-f628-4ca2-9ebb-4d6f608633bf · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

TopK Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:41.035447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:37.854065Z digest=sha256:7129fb36421ce48d6de30c325713e5049e8f4bbcba8a5e6744713ba66062de1e

Observation d02c0341-8289-4f6a-9e93-7c948279a759 · outbound

This paper cites Improving dictionary learning with gated sparse autoencoders, 2024.

TopK Language Models Improving dictionary learning with gated sparse autoencoders, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:37.972247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:37.972247Z digest=sha256:272fab14f51b5df85acb679a460ba0cc98a30f976afd2d42b44208bd1005c8f6

Observation acedd645-2f42-434a-8ef1-aebff0af90bb · outbound

This paper cites Winogrande: an adversarial winograd schema challenge at scale.

TopK Language Models Winogrande: an adversarial winograd schema challenge at scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.142907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.142907Z digest=sha256:60469d9ae735786741541dfd7a75064eaa56c49f7d0f73af86e7b92ad4236acc

Observation 8f6dab73-abb4-4500-b8d5-f326759064a8 · outbound

This paper cites Taking features out of superposition with sparse autoencoders.

TopK Language Models Taking features out of superposition with sparse autoencoders

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.851527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.258973Z digest=sha256:7a4fe0c60afe72261e4a7d0a07a0286ef7ba1dc3097988b177a43439ddbb4d9d

Observation ed8b5f35-897e-4c8d-9f38-e00bfacbbac6 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017.

TopK Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer, 2017

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.370614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.370614Z digest=sha256:c372ae2249ccb23ca883fa8b1cbf720d19f0eedcbd250d8a89d2ecf96e9f5e2d

Observation 1ddad81e-ea7a-41e9-b420-c9e5d64197bd · outbound

This paper cites A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025.

TopK Language Models A survey on sparse autoencoders: Interpreting the internal mechanisms of large language models, 2025

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.670077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.509917Z digest=sha256:ed36bc1d44af88c3dd79f0de7116b578e49d59d594577ce5d48a2c6cf56ed479

Observation 9c0496bf-71fc-42ea-a76c-b215fc6f7781 · outbound

This paper cites Codebook Features: Sparse and Discrete Interpretability for Neural Networks.

TopK Language Models Codebook Features: Sparse and Discrete Interpretability for Neural Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.680344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.680344Z digest=sha256:32953413ab435a0e2b99e50115df8bb2774225a54cad6b7334152acab16d5f98

Observation e1b8015f-8dca-48ee-a6d5-94e0197c7317 · outbound

This paper cites Daniel Freeman, Theodore R.

TopK Language Models Daniel Freeman, Theodore R

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.462326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.794923Z digest=sha256:b030d17ab211aeb47e1cd034e95c98ebce1dcff4d25d96f72907e70b10b4088c

Observation 4e902680-9768-47c0-aae7-a1a9b5f98c7c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TopK Language Models LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:38.890079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:38.890079Z digest=sha256:2cfff9be2aedaf97dcd0dcd44ba17e5a13480b025f7a3b00aa6c9ca574ad6ce0

Observation 3ea79e63-1c97-4e4a-ad6c-80dc996c16ab · outbound

This paper cites Meta Lingua: A minimal PyTorch LLM training library, 2024.

TopK Language Models Meta Lingua: A minimal PyTorch LLM training library, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.282755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:38.998159Z digest=sha256:fec8ad45cac805702e0cd91cef1ae0cdbcfa9ecd0d3a429e0af1e073e4730955

Observation 373449a0-f1a3-45e0-b63a-24c6155ab499 · outbound

This paper cites Tracking the feature dynamics in llm training: A mechanistic study, 2025.

TopK Language Models Tracking the feature dynamics in llm training: A mechanistic study, 2025

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:40.093824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.060459Z digest=sha256:c5833d4fdeca7b55863a485165748a5e129ddfb65466ce7b732e3040210801ef

Observation b26d6674-e1c9-49fd-a3d2-63892221c21a · outbound

This paper cites STEP: Staged parameter-efficient pre-training for large language models.

TopK Language Models STEP: Staged parameter-efficient pre-training for large language models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.921354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.121952Z digest=sha256:562fec5d8ba6be6a515de7a82e5c25b27bd7d3b07a50086899da1e2d5bb4d6e5

Observation 4a13895a-5f6b-42e3-a2c2-6056c77dd21a · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019.

TopK Language Models HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics , pages 4791–4800, 2019

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:31:39.727732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T22:31:39.176817Z digest=sha256:486efc07918df487e04c17ab91246e639c0e5d48214fb2c874e2d7fcf596bcdc

Observation 757b2521-884d-4004-a072-b40518b8c712 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TopK Language Models PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T22:31:39.220992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.220992Z digest=sha256:adbe63affa5cd918379fdff6d42e07d0f4e215d90646e31aa4befb0ec1b3918e

Pith citing papers

No inbound Pith citation observations are available.