Pith. sign in

Paper Citation Record · LEDGER

An Analysis for Reasoning Bias of Language Models with Small Initialization

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2502.04375.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04375 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:28:52.279862Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:15.054309Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:05:04.830513Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e863a94a-9e15-4fd7-a0af-a7c4fd290b6a · outbound

This paper cites Phi-4 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.220951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.220951Z digest=sha256:35c0316e2abc3eccb51c912bcfd88a4c8c18fd13f672eaf27914968cc25b9fca

Observation 572a90c5-ba55-436e-9ae3-e26e705ce555 · outbound

This paper cites GPT-4 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.227174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.227174Z digest=sha256:56219e30bd95a9a9ad3984057b9bc3829be437efeb5f4cb219319cb2e441335d

Observation 7612f387-27eb-47fa-9cf5-e1ead065d700 · outbound

This paper cites S., Hu, W., Li, Z., Salakhutdinov, R.

An Analysis for Reasoning Bias of Language Models with Small Initialization S., Hu, W., Li, Z., Salakhutdinov, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.666924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.232212Z digest=sha256:a6ec6245d1248ad8aa0de1c54b4193dab4d1f33237b851200eadb240b13c7845

Observation bb436b54-89ab-415d-bcfc-f53260046733 · outbound

This paper cites Reflections after refereeing papers for nips.

An Analysis for Reasoning Bias of Language Models with Small Initialization Reflections after refereeing papers for nips

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.616478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.236969Z digest=sha256:af10d012735d87aadf8fa22910a94b890ff501c05285970f318a680267695978

Observation 01915d98-d04e-4ebe-9319-62ae92630134 · outbound

This paper cites and Rathie, P.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Rathie, P

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.525810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.241569Z digest=sha256:b9c65eaa92c89a6207fce532991b3777b6ee0e2e0abb30d4230bad7e95ddd050

Observation 7280dc0a-53e8-4036-8f73-99b8c833acf1 · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-09T05:28:52.349960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.246643Z digest=sha256:4f3a7dd32f79964e3ad24cf8c8e3c3e4cf97658948f573852af91ad8c4e8329e

Observation 0d67c68a-3e61-4241-829f-9a3692893d1b · outbound

This paper cites and Bach, F.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Bach, F

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.464356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.252040Z digest=sha256:23f9a43a2a25c8e4cc8eb2a1460e5bdc791ae04edf342ffb457cdfac8a13cbe1

Observation 2ed755b8-7569-4ff0-89d1-601a80a9f9b5 · outbound

This paper cites Faithful Reasoning Using Large Language Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Faithful Reasoning Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.257146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.257146Z digest=sha256:255e0da754a3c3dc3f2e6e91c8b2413b6b65ec78abb163cbaccec146622d6084

Observation 54507407-694a-4d15-9c72-0177863cc7de · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

An Analysis for Reasoning Bias of Language Models with Small Initialization Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.262280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.262280Z digest=sha256:13f10a039758d2039f25a37972bf856575227685ed6369bea49873693ccaeb16

Observation ece0bfbf-5f29-4af2-8523-40b8fc116d48 · outbound

This paper cites The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization.

An Analysis for Reasoning Bias of Language Models with Small Initialization The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.278903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.278903Z digest=sha256:7b156b8190f870ea4f6cfbc99ba1eb65f70485f21f8313745f4713cccf6fc77d

Observation 36299e03-8cb4-4498-8d23-30005147fb43 · outbound

This paper cites CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural Representations.

An Analysis for Reasoning Bias of Language Models with Small Initialization CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.310278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.310278Z digest=sha256:0a6b71d0d683b80aa2f43a74c9cd18a824a36e68fb9eb468ee4bacb8cb59d6a1

Observation 8344b049-1e7f-4a95-9e9d-78aca8964860 · outbound

This paper cites L., Jiang, L., Lin, B.

An Analysis for Reasoning Bias of Language Models with Small Initialization L., Jiang, L., Lin, B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.343200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.343200Z digest=sha256:5b809abc55a1936e8e193e57d356644734614986e6a4577d48ed442679382f7f

Observation 5c83dc4b-a778-4551-9439-c7bc29679656 · outbound

This paper cites A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics.

An Analysis for Reasoning Bias of Language Models with Small Initialization A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.422368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.375582Z digest=sha256:f5f91e8f994e5cc71e1621b4753c6e09b19f31c09f5c48db3af7c5343159f9e7

Observation 73bb8b33-8326-404f-a4e5-7a52f025a538 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

An Analysis for Reasoning Bias of Language Models with Small Initialization TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.403494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.403494Z digest=sha256:1232bbdc6045e1bc977d612ce21dfe8936efca5dd0a606735ba67d35be81e96c

Observation f571c1fe-67d4-478d-84f4-639840cf1a11 · outbound

This paper cites How does gpt obtain its ability? tracing emergent abilities of language models to their sources.

An Analysis for Reasoning Bias of Language Models with Small Initialization How does gpt obtain its ability? tracing emergent abilities of language models to their sources

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.406308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.442629Z digest=sha256:8bd9b7e80f4ecf886a1c1e5580b3aff93684eaf9114587feb6dc8da9042601b8

Observation 671aca84-d7a9-4b9b-b25c-0c7a3538cda2 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

An Analysis for Reasoning Bias of Language Models with Small Initialization Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.501648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.501648Z digest=sha256:ecbe8593fa94148dbf9628efb1a2651039eb6a91d8d84c28953f6a45ac0767e0

Observation bce18684-9761-4eff-b1b7-d8fc826d7f09 · outbound

This paper cites S., Perez, F., Ba, J., and Volkovs, M.

An Analysis for Reasoning Bias of Language Models with Small Initialization S., Perez, F., Ba, J., and Volkovs, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.379227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.547458Z digest=sha256:f2ff2a520a37659ea8fb60b377ca5b0dd1a10235a4b00743d458a813a7baa189

Observation eefeb4ba-6f10-4d26-85c3-dcebc66f833d · outbound

This paper cites Learning compositionally through attentive guidance.

An Analysis for Reasoning Bias of Language Models with Small Initialization Learning compositionally through attentive guidance

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:28:52.745846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.552290Z digest=sha256:9625c51d0f9d8c09f5d8926432f828b390acdc0c2b5b776ff76db21b250b6b9c

Observation 0fa9d1e1-9024-4aa2-bee3-4c2504ddd66c · outbound

This paper cites Neural Tangent Kernel : Convergence and Generalization in Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Neural Tangent Kernel : Convergence and Generalization in Neural Networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.363554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.557359Z digest=sha256:8e25a4b0a94e6273cff2f49140041c75e172124fb3228a143966b44bc8404c7d

Observation 718e5914-bfe6-4e6b-88e0-52a57fd03c0d · outbound

This paper cites B., and M \"u ller, K.

An Analysis for Reasoning Bias of Language Models with Small Initialization B., and M \"u ller, K

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.562035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.562035Z digest=sha256:bf3a7ae60e9d78dcdbf2095ee53cf0e5c27686874612c893a3d7c2679aa93e82

Observation d68ae65b-4b0e-4cc9-9b77-c09de6140f3c · outbound

This paper cites Break It Down: Evidence for Structural Compositionality in Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Break It Down: Evidence for Structural Compositionality in Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.566749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.566749Z digest=sha256:b10190c881821a0a65081667fb0ede2d0f62151d9a3eb7666ceef198a8967595

Observation f9f4ffcc-a086-4b30-9e69-316ee80546cf · outbound

This paper cites Not all tokens are what you need for pretraining.

An Analysis for Reasoning Bias of Language Models with Small Initialization Not all tokens are what you need for pretraining

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.571389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.571389Z digest=sha256:25b8b24893bc510edd6df5c0b4f929177d37b0d7ab0a85586ec5708d84db2055

Observation 7d20637e-1d00-4009-a74c-95afaac33a48 · outbound

This paper cites DeepSeek-V3 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.575637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.575637Z digest=sha256:a346f3c4a5cdfc09c8f14a7871e3791bb29ed7c01f46bb3c4e27f9ae0b7c632d

Observation 276bfca0-e664-4254-95a6-d80ad59018ca · outbound

This paper cites Transformers Learn Shortcuts to Automata.

An Analysis for Reasoning Bias of Language Models with Small Initialization Transformers Learn Shortcuts to Automata

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.580568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.580568Z digest=sha256:598700a10faddc9e5a9eded838c762ac0207bb1718355e4348650768415d74e2

Observation a35bad4b-12e6-4bb5-9f02-49f411c4d22c · outbound

This paper cites Understanding the Difficulty of Training Transformers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Difficulty of Training Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.585595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.585595Z digest=sha256:7eeb428726118bb6214126b70446fe3ef7ab12be7e1b5d069526760e75261958

Observation 17c8b97d-acaf-48b7-b6b2-195ea4becdf4 · outbound

This paper cites J., Ma, Z., and Zhang, Y.

An Analysis for Reasoning Bias of Language Models with Small Initialization J., Ma, Z., and Zhang, Y

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.319975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.589800Z digest=sha256:3ada451d638425a4d006ea67128772e5133ce503716eb042a31b6b4bdb836491

Observation c396f774-5388-4fc7-85ab-05341015456f · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.594177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.594177Z digest=sha256:8a04a73fe04859860aafb6fa2ef87a0b83ab8e30ec33052ecf649e0b3295ef60

Observation 318fdf09-8ad9-422b-8499-08325e10cf91 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization A mean field view of the landscape of two-layer neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.598613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.598613Z digest=sha256:40f623addfb92969821b36ba163d86e991ab3dd377e80fbdc4ce6326638744c3

Observation e96eff98-44a6-4c7b-af36-3131d9171e0c · outbound

This paper cites Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task.

An Analysis for Reasoning Bias of Language Models with Small Initialization Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.603410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.603410Z digest=sha256:89797decb2ac27548cb3197e93752cb4960f8fa475966123257894a18c763056

Observation 18875ebd-5a96-486f-883c-61a8d0949742 · outbound

This paper cites Language models are unsupervised multitask learners.

An Analysis for Reasoning Bias of Language Models with Small Initialization Language models are unsupervised multitask learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.620847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.620847Z digest=sha256:8c3da68cbcf916ac769541831de1948e5d9fd78e57d74520c1695bee9aa8044a

Observation 857d318c-56eb-4190-b762-acc542225018 · outbound

This paper cites Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.641065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.641065Z digest=sha256:a254ce860c391297dc8366d5a61473b3051eee5ec417f0bced0365eddbb59776

Observation e81b832a-d782-4d54-97f1-9355dea9d994 · outbound

This paper cites and Vanden-Eijnden, E.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Vanden-Eijnden, E

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.285195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.700823Z digest=sha256:9aa0f43849a3810c6daac0a1a2fc6b94aa0bf4b679ac732db435ea46b745d0ad

Observation 82d661f1-f854-4d06-8bfb-bd376c14dbac · outbound

This paper cites and He, H.

An Analysis for Reasoning Bias of Language Models with Small Initialization and He, H

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.269539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.738043Z digest=sha256:ff50add707aedec8c62165c3d695660aa0762c24a381abd0797effce1034f7bb

Observation 485a5d23-6b5d-4d98-8247-abf04b726b2d · outbound

This paper cites and Spiliopoulos, K.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Spiliopoulos, K

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.757701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.757701Z digest=sha256:47b05883514d2e9d9dc53c92b2109110ba16ecc59c5ac26431f08ba0311a4887

Observation 6756689a-d32d-49bd-9937-a4f787d68eb8 · outbound

This paper cites Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems.

An Analysis for Reasoning Bias of Language Models with Small Initialization Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.255005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.784618Z digest=sha256:ede0dcf7cd9ce861f8b9b722bdb4bf311b265208101f3981d3a2eb7d75089a96

Observation 42c48d34-487b-4c4c-a673-849b591f0d9d · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.788616Z digest=sha256:50ea7c646880c7b79c5033a97560d18f094d41c8b503dc328d516438dadbc48f

Observation 58695419-ce1f-4005-9824-f486b8936afa · outbound

This paper cites and Kolter, J.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Kolter, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.239723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.793071Z digest=sha256:2a8a3bc4ca9cee9a9c30db326c3d8ae5bb6d72087e66deedb61205c0d5da228a

Observation 4569c100-5d1e-4abf-bf1c-3329d336f44d · outbound

This paper cites Deepnet: Scaling transformers to 1,000 layers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Deepnet: Scaling transformers to 1,000 layers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.224112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.797195Z digest=sha256:f9df8f9f3d1bc7dacd129fd2c54492f0abf71e26ba4f53d5f797726585da425b

Observation bd5896f8-8a41-4384-8067-67e48f6e9198 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

An Analysis for Reasoning Bias of Language Models with Small Initialization Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.801401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.801401Z digest=sha256:fc1c42fa5a8f3f8effbe6b39faf5fee3dc3400e329ecfa34086726ac3da8902f

Observation be13f3e7-9ac0-4b58-9233-54524425ebdd · outbound

This paper cites Improving Generalization and Convergence by Enhancing Implicit Regularization.

An Analysis for Reasoning Bias of Language Models with Small Initialization Improving Generalization and Convergence by Enhancing Implicit Regularization

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:28:52.592722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.806335Z digest=sha256:dfc9d53d599a93ed4044f98eca2fa9fecdfd0ef946a860709a9ac26c65d51d5e

Observation ed352f33-948f-4aad-a02e-453e9df0eec5 · outbound

This paper cites Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.810922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.810922Z digest=sha256:db322f7e231ae386dccc989ec90eeb65b3e6a7376e94f607d456d29de1eea54d

Observation 39617b73-84ee-4467-93fb-da96bb31208f · outbound

This paper cites Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.815997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.815997Z digest=sha256:138b432bbc18922485b2eb8928895aa89206d1cd07bff1e74b74dce18ec76e0a

Observation cd478da5-6f8d-4d05-b085-6e5dad8133da · outbound

This paper cites H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W.

An Analysis for Reasoning Bias of Language Models with Small Initialization H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.199486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.820681Z digest=sha256:80c4c683b6dadd34eabb3e21d08b779fa2c8ab4606cd62e3a56bfa9e34bd1f4b

Observation b2f1df8b-9f3c-4c1d-b4b7-40cc19ef6a2f · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.824742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.824742Z digest=sha256:1b9df47687390b6e69c98f84b9f39263c223f4198ca77e469729b0a37b7890b8

Observation fb5e145f-64c4-4f27-bd06-d99b826300bc · outbound

This paper cites Gradient Dynamics of Shallow Univariate ReLU Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Gradient Dynamics of Shallow Univariate ReLU Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.829313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.829313Z digest=sha256:3e2693eb7f6f06958eba470f37c6dfc2575647176670d51727f2b19972cf0657

Observation b8ea9d09-8e59-40d0-912b-10b067bfa388 · outbound

This paper cites An overview of condensation phenomenon in deep learning.

An Analysis for Reasoning Bias of Language Models with Small Initialization An overview of condensation phenomenon in deep learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.834240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.834240Z digest=sha256:cb412fc175051772bf8e3267cdc062e4456c5d6a19079b2424992fca9aa3dd61

Observation 54c580f6-6c77-4e44-9636-09a75b1d5d17 · outbound

This paper cites Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?.

An Analysis for Reasoning Bias of Language Models with Small Initialization Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.838929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.838929Z digest=sha256:e7b4d613f89997b541f55396f551c8cec2a2bd284a075ffe769e3a2e1e7fd48f

Observation 172990d3-9470-4d19-b1b3-d190604287c2 · outbound

This paper cites Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention.

An Analysis for Reasoning Bias of Language Models with Small Initialization Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.843650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.843650Z digest=sha256:95e8770e2d0c905df19a8bcef2b1cb9953ff094af02cb3dce0a3fa0359d5e5a0

Observation fe9d802e-28cb-41b2-a937-ba971030ca36 · outbound

This paper cites Understanding deep learning requires rethinking generalization.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding deep learning requires rethinking generalization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.848843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.848843Z digest=sha256:dbb7a13abcc53ac21b9fb3932108fb58534d4a0ccbc8add8eda4b51a007a31f9

Observation 38bcef4a-6788-4061-a647-f77c37aae9d4 · outbound

This paper cites A type of generalization error induced by initialization in deep neural networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization A type of generalization error induced by initialization in deep neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.853798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.853798Z digest=sha256:8daee66eb41644bc3d2af19ce48642c6032e010c9122a9faf6abf36f8de49f69

Observation add21f00-2939-48fa-90f9-e7ad69b09681 · outbound

This paper cites Linear Stability Hypothesis and Rank Stratification for Nonlinear Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Linear Stability Hypothesis and Rank Stratification for Nonlinear Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.858900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.858900Z digest=sha256:670c0c0915686cde62925883ba8d8ea32f491698388d6b80a9af89979b092570

Observation f822a794-272e-45e8-9b20-e0794a811df3 · outbound

This paper cites Loss Spike in Training Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Loss Spike in Training Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.863442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.863442Z digest=sha256:a77a85c214e840b1501ffa06a0cbaeaeba31c16e5e64355b82c70e564e323540

Observation 1c412bb4-6094-4512-b027-703e74fbdc37 · outbound

This paper cites and Xu, Z.-Q.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Xu, Z.-Q

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.141956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.928943Z digest=sha256:a1452d2af11cca009b7a529c2613e1230f447be9c9c68ad52878e294826b1eec

Observation a52381c9-da47-499e-8db9-e5007416eeac · outbound

This paper cites Stochastic Modified Equations and Dynamics of Dropout Algorithm.

An Analysis for Reasoning Bias of Language Models with Small Initialization Stochastic Modified Equations and Dynamics of Dropout Algorithm

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.046608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.046608Z digest=sha256:5ed3422f20a91afd200dd3b9bb18fd7ab1fb61d3ed7b83f627a39745f2a61262

Observation ec5d6f10-a553-4fa8-b59c-f0dcbcacc58d · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:28:53.029678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.153482Z digest=sha256:2ad26c5c4839da09caa2942aacd1255d4677ba5978a779118340d2dcafc255a9

Observation ef590fc0-5244-4723-aa78-0af3468e01e1 · outbound

This paper cites Anchor function: a type of benchmark functions for studying language models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Anchor function: a type of benchmark functions for studying language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.234294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.234294Z digest=sha256:b132dd8b7b9a01ff7dd4e7f27792afca8e7f080dc502babf4fa9e4d4bc43a2bc

Observation f2c555c6-f9b4-4532-8140-f292ba9d8258 · outbound

This paper cites Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.265267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.265267Z digest=sha256:8cdbce2a73983cedd3540fbfd872c6d56318710bace49ac794b2eb8824e34a1d

Observation d9937e27-fdfb-49b8-bc92-a1b7cf3b48f5 · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:28:52.931262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.270541Z digest=sha256:f998cc21ea3db45df1a4876324906bbfc9ddb26bcb2b2e3ee5b2e5b88aeaed9b

Observation 0c30e867-4b53-4098-b86a-c8789600bfe7 · outbound

This paper cites R., and Goldstein, T.

An Analysis for Reasoning Bias of Language Models with Small Initialization R., and Goldstein, T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:52.897390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.275817Z digest=sha256:97dbcc65d898e1b9132d4889fcfbfba9e38a42c612f95157b80925163d2200ab

Observation 2bf11d20-18e9-4d4b-9d59-2321367ceb6a · outbound

This paper cites write newline.

An Analysis for Reasoning Bias of Language Models with Small Initialization write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.279862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.279862Z digest=sha256:fac7781001f5e332dfabed4f7f86b74e9c57505eab5d2eb417d71e320308aa2f

Pith citing papers

Observation 66fefe0f-4f93-47e8-9afe-9276c46157ab · inbound

An overview of condensation phenomenon in deep learning cites this paper.

An overview of condensation phenomenon in deep learning An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:05:04.833173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T20:02:41.105786Z digest=sha256:def104357c077f88308cc43ffe33551b86e4fd57359be7cddbb96ac21a2cf9f1

Observation b8fc0767-8d84-497a-abfa-5756056ee175 · inbound

Scalable Complexity Control Facilitates Reasoning Ability of LLMs cites this paper.

Scalable Complexity Control Facilitates Reasoning Ability of LLMs An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:15.054309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:15.054309Z digest=sha256:81b992633eed38715e7ca4484c4abb67cdaed3130f31af4b88c54bf428a91b04

Observation 6ec5b63b-4bbc-4e57-864d-1447d941c7f6 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:18.723775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:18.723775Z digest=sha256:b2cb31804161944ea0b419d669f8b53acaac501434bc2791b8eb5deb1b80a1f1

Observation a616f226-79c2-4b59-a397-8de0f7fbb763 · inbound

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability cites this paper.

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:52.679420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-16T11:20:33.400885Z digest=sha256:e778bb87a54a99ec8c49f86cee9ef3df749875bc45bc90c1ef90f2eb4b025cb2

Observation 1c17a3da-46b0-4515-aaf6-d2123e7139f9 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.250276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T18:09:54.899039Z digest=sha256:f3b60cd2b2fff0981daa6e04218e62d20df2a821958d82309f63f6f2819f79f4

Observation 978df83f-d197-48a9-84d2-e4b6899c320e · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:49:44.634914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:49:44.634914Z digest=sha256:20467910f4503bd353652885feac84990093a1957b2732762a4f5386dd03dbd1