Pith. sign in

Paper Citation Record · LEDGER

Reasoning Bias of Next Token Prediction Training

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.02007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02007 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:46:43.619470Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:50:43.217179Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact5
  • verified fuzzy23
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 908f4876-0b39-4468-b68b-25e6da6230ad · outbound

This paper cites Phi-4 Technical Report.

Reasoning Bias of Next Token Prediction Training Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.282422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.282422Z digest=sha256:4be4366a8890ff951f002cc3dbf6054faa5369f5d6f3c91528d0a745f9b2dc45

Observation 0bc3db89-9f45-4529-a4eb-512b5c8a79a6 · outbound

This paper cites and Nagarajan, V.

Reasoning Bias of Next Token Prediction Training and Nagarajan, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.818716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.288710Z digest=sha256:007a9f615a62ffa829bd0882ea713f4660d3fbd22aa33c377b192dfdcbf2280a

Observation 5a091145-20fd-4987-a9f7-37a6a77efe5d · outbound

This paper cites and Giryes, R.

Reasoning Bias of Next Token Prediction Training and Giryes, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.803082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.294502Z digest=sha256:c6583473c2a7947c87e92d1c0dd1efb017d12a7caca0e7c5c9545293d2ad4ba6

Observation 54745b23-35fb-4dec-9396-4dc895cbfc58 · outbound

This paper cites Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation.

Reasoning Bias of Next Token Prediction Training Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.299577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.299577Z digest=sha256:2b3d39f74032cbaaf347b6cdb16b3ddb6c1420a6375ccd5d5048392d37c293bc

Observation b0fb012e-f872-44c1-b650-21a26278852c · outbound

This paper cites Understanding robustness of transformers for image classification.

Reasoning Bias of Next Token Prediction Training Understanding robustness of transformers for image classification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.786239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.305137Z digest=sha256:1e4fa0b45bad5db49145f8c9c333d50606b21a1cd48073d2b726d1797239474c

Observation b1e68e2b-b26c-43cb-b507-2af169b07173 · outbound

This paper cites R., Angeli, G., Potts, C., and Manning, C.

Reasoning Bias of Next Token Prediction Training R., Angeli, G., Potts, C., and Manning, C

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.310512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.310512Z digest=sha256:e4f2c4dce83d824ad86601985b453cb49d4be4ea24d6c8b174cb1c85fd5d5c49

Observation 64acb6d5-312d-4437-a06c-2ecbe4500664 · outbound

This paper cites Language Models are Few-Shot Learners.

Reasoning Bias of Next Token Prediction Training Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.316250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.316250Z digest=sha256:0038a7c5f727779ea9d04376bfe5e77d7f5f54cd07c6182c9b536a4d1da6727c

Observation 26d82bc1-9b13-477a-b139-85f4720360ab · outbound

This paper cites Dropout as a low-rank regularizer for matrix factorization.

Reasoning Bias of Next Token Prediction Training Dropout as a low-rank regularizer for matrix factorization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.770532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.321652Z digest=sha256:72ca7bb5a933c858bb177bd68bde4cf865a71d510224a026172f655e04ad3ca0

Observation 72c42146-f7ca-464b-8eeb-13a91c8e4e1c · outbound

This paper cites Transformers as soft reasoners over language.

Reasoning Bias of Next Token Prediction Training Transformers as soft reasoners over language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.326651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.326651Z digest=sha256:47f1a05ea5c8f3a6392bfe400892e1fe96e06dedcb62363f9519b67b4414c296

Observation e1d7fe60-f30a-416f-a531-3f5520c0b2ca · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Reasoning Bias of Next Token Prediction Training From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.331535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.331535Z digest=sha256:fbaacb006de9a378a58ad12562631aaa6a2ac8312b9d7b12bb3b39b98b710118

Observation ec8fa746-963f-4ee2-a080-7addc029617f · outbound

This paper cites Reducing Transformer Depth on Demand with Structured Dropout.

Reasoning Bias of Next Token Prediction Training Reducing Transformer Depth on Demand with Structured Dropout

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.336880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.336880Z digest=sha256:826fd04b468a3cc06939e0f259fa53ce54f09c8e88f0d7c67210e11f8cd96025

Observation b3ed6204-64b0-475a-99a1-9b58e512403f · outbound

This paper cites and Tu, Y.

Reasoning Bias of Next Token Prediction Training and Tu, Y

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.754241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.342409Z digest=sha256:32b90584ed2cc39e4c364faadcad300b2de6e12130ba0a1c8d245d132cd9bc6e

Observation 004c79a9-e648-4a5a-9026-951e34de47d9 · outbound

This paper cites Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G.

Reasoning Bias of Next Token Prediction Training Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.739085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.347125Z digest=sha256:202f87b05645d8f2ce920b2483c87876a787448f5684ae789b7607349b474f57

Observation dcf418dc-aec7-462a-a395-17c6feb2b8d6 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Reasoning Bias of Next Token Prediction Training Training Large Language Models to Reason in a Continuous Latent Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.351846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.351846Z digest=sha256:3fcc0b406836fe445704aff1eda11b5f68405822ba0cb47578fef4734b8932ed

Observation 0373d6dc-8f9b-411e-b6a5-e400175c010a · outbound

This paper cites A Law of Next-Token Prediction in Large Language Models.

Reasoning Bias of Next Token Prediction Training A Law of Next-Token Prediction in Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.306519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.357172Z digest=sha256:732f7df2aa8d3ea9adfd9042431430b2427c10d311983b200a3ad7fd800c5410

Observation 5889224b-91c3-4f62-977f-ef673b0eab86 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

Reasoning Bias of Next Token Prediction Training What Matters in Transformers? Not All Attention is Needed

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.362438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.362438Z digest=sha256:4ae90f32ab325b43a49796e57d6b3a25ac725e58f68bda024650c402ec257f4a

Observation 42335d35-d8bc-4683-8f70-55e86a492a40 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Reasoning Bias of Next Token Prediction Training Pretrained transformers improve out-of-distribution robustness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.367522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.367522Z digest=sha256:993ff5a6d806b0b4a908529cc4491e3fce52fc12a91692dff37fd8a6986d2812

Observation 80767613-99e7-4bfa-9773-0c646eae34da · outbound

This paper cites and Schmidhuber, J.

Reasoning Bias of Next Token Prediction Training and Schmidhuber, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.724058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.372397Z digest=sha256:22a8b8c655bcca7ac24e235f906f14afcb42d655ead93427d7ccd7983f2f9ec6

Observation d3608fce-dd3f-4a81-94ca-1d129db31b3e · outbound

This paper cites NEFTune: Noisy Embeddings Improve Instruction Finetuning.

Reasoning Bias of Next Token Prediction Training NEFTune: Noisy Embeddings Improve Instruction Finetuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.378061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.378061Z digest=sha256:388ccb1fe5ccad778c71573efec21faab16db825b53296d5a3f6c6e2362549e4

Observation fdeedeed-ec45-46a3-90c1-5753574ac00b · outbound

This paper cites S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P.

Reasoning Bias of Next Token Prediction Training S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.383548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.383548Z digest=sha256:bcdcf923fc3e726c97c2cbf1fe865ebce9130bb458599e71fef63465b0757cd9

Observation d34e9dcb-e86e-48a9-8da7-8b7bdf6dbfce · outbound

This paper cites N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al.

Reasoning Bias of Next Token Prediction Training N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.697997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.388908Z digest=sha256:8b8f4a6e7645d002e94947bf7b05b625e4cef7b083b28a7d39b14bbf6ce14ebf

Observation a0bf94c7-c237-48c6-9fb2-05d1b59d2d50 · outbound

This paper cites J., Xing, E., and Caruana, R.

Reasoning Bias of Next Token Prediction Training J., Xing, E., and Caruana, R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.681991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.393776Z digest=sha256:02bd421feaa94ee8bd8e54dcabfddd7ec7a86b2913ffe0a39871d7310f78b137

Observation 599dc58c-32e0-4abc-9f01-e43c660afdad · outbound

This paper cites DropKey.

Reasoning Bias of Next Token Prediction Training DropKey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.398762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.398762Z digest=sha256:c5638e09841215c5034217fbc28fcf997bf4f0375c4732f9623876749d806c6e

Observation 045c0bc8-f000-49af-bbff-6e95ce11acf6 · outbound

This paper cites Challenging large language models with new tasks: A study on their adaptability and robustness.

Reasoning Bias of Next Token Prediction Training Challenging large language models with new tasks: A study on their adaptability and robustness

Reference 24

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.800168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.403768Z digest=sha256:8489cfa14d8e5d01c05151344e12d8183867719642bf4d2d159fa138a8bf8fc9

Observation ce9ae42e-ea01-4df8-91f7-653a262e3cbd · outbound

This paper cites Visualizing the loss landscape of neural nets.

Reasoning Bias of Next Token Prediction Training Visualizing the loss landscape of neural nets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.665175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.408608Z digest=sha256:ee555b905de8bb5e9d8fb57efc90657847f23742cb7706485fa8b5fa7dca0397

Observation 71f357b0-9897-4fb4-9bc3-24524f166e75 · outbound

This paper cites E., Singh Rawat, A., and Oymak, S.

Reasoning Bias of Next Token Prediction Training E., Singh Rawat, A., and Oymak, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.648043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.413206Z digest=sha256:52a1fedf126ce3bac4a647390bff211885880a05b58b45bb7f0168ac1f69bb8f

Observation 638fbb91-2244-4cfb-bbf0-bd7c71d5aa05 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Reasoning Bias of Next Token Prediction Training Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.418162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.418162Z digest=sha256:b078077de0180cc561491cda1e3e2d38736ec7a697d43cd212a5cc598b505f45

Observation bb035bc5-c676-4945-be4e-7ed6c8263df2 · outbound

This paper cites M., Li, Z., and Ma, T.

Reasoning Bias of Next Token Prediction Training M., Li, Z., and Ma, T

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.632706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.423209Z digest=sha256:28e2f2eb72d51cea59d4d80ebfdbf41498b3f30a9f11e5e245c991df70f0abf8

Observation 405bb9c5-4ebb-4351-829a-b8ec1772e82f · outbound

This paper cites and Ying, L.

Reasoning Bias of Next Token Prediction Training and Ying, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.616824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.428213Z digest=sha256:f043cca397bc3e9b0a294c4b53859752fdc6760a4e9e80b6b731a65c9e7ad067

Observation 2ea55aa7-4095-49a2-8ac6-d49fba158f4d · outbound

This paper cites Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024.

Reasoning Bias of Next Token Prediction Training Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.433208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.433208Z digest=sha256:8c4c1b56c64f78fb962bc87bf2eef7d325a1a42bf23166476d9ff7dae5975931

Observation 3f18b2ec-3d31-4193-9860-5a52fe1ae1bb · outbound

This paper cites On the implicit bias of dropout.

Reasoning Bias of Next Token Prediction Training On the implicit bias of dropout

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.601307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.438140Z digest=sha256:d3481dd8ca5d1616a6f9eb572eb5d1fb34c59432b92d0d8c740c8e39f3f9a86b

Observation 53d00bd4-9b33-4657-b951-5a431f8cafe1 · outbound

This paper cites Pretrained Transformers Do not Always Improve Robustness.

Reasoning Bias of Next Token Prediction Training Pretrained Transformers Do not Always Improve Robustness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.127239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.443060Z digest=sha256:fd40117fae40d4f93161757515102b0223322d59ff7ad988da6082133c533fd2

Observation 128a33fd-553f-42f0-a89e-e447c3a3bb11 · outbound

This paper cites and Samwald, M.

Reasoning Bias of Next Token Prediction Training and Samwald, M

Reference 33

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.783504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.448286Z digest=sha256:94b60a57d8b74f283dc34d0c22d7a5afd2d1f28ecd22a616cc6741af7bf986eb

Observation 79db3ce0-125d-490e-9828-df83de123781 · outbound

This paper cites Power-law escape rate of SGD.

Reasoning Bias of Next Token Prediction Training Power-law escape rate of SGD

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.453263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.453263Z digest=sha256:582b664ebdd39efec4db2e02e5d96f3033e904d007485cdb652dfe9e34ecfcf1

Observation 36c2b9ac-3e90-46a9-8e47-10777162b5a6 · outbound

This paper cites LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models.

Reasoning Bias of Next Token Prediction Training LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.459652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.459652Z digest=sha256:09fa6cc0f0e4280f1cab955e2805125160b22f7c774fc469d1cb53168f9bcae6

Observation 991996d3-8a7c-4857-b014-46ed8f8781cc · outbound

This paper cites and Narasimhan, K.

Reasoning Bias of Next Token Prediction Training and Narasimhan, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.584285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.464879Z digest=sha256:42904e5b4a232e1289cf1797d0a4fecdfb6b51461ea70901adbd210f254eecd7

Observation 7cbfafc1-b7aa-4038-b36f-4521b0666e6d · outbound

This paper cites Language models are unsupervised multitask learners.

Reasoning Bias of Next Token Prediction Training Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.469944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.469944Z digest=sha256:55e15ff514da4e144594ac4e4da256e2ee4449c43fbf1f4b735a555a801a45d7

Observation d24bcdc6-d40d-449b-b2bd-fdd6d1898055 · outbound

This paper cites R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners.

Reasoning Bias of Next Token Prediction Training R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.474845Z digest=sha256:d32a221fb863d0369c00eac11dc3aaa9677e11d99cf9e3821c6693ece996847a

Observation 1242e8b6-90c7-41f1-a5a8-e94f5255b8ed · outbound

This paper cites and He, H.

Reasoning Bias of Next Token Prediction Training and He, H

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.479930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.479930Z digest=sha256:fb2544c9a3e1ccdbbce33713800edacb6681fa23fec54bc6f57cdd7d9f107f12

Observation e0def47a-ad1f-4840-9711-164b47e0c829 · outbound

This paper cites Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts.

Reasoning Bias of Next Token Prediction Training Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.485090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.485090Z digest=sha256:a39a3081e32b9189858472e4610f79605959fe7a385a8b72f089b7b4a53f2d5d

Observation 03ca9248-61bf-40b2-8d41-56918fad7417 · outbound

This paper cites an unresolved cited work.

Reasoning Bias of Next Token Prediction Training Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.490010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.490010Z digest=sha256:968110c22d0ad600964ef00f5bdb8e74d80d7fb90d267a7925cd7832ec16ff41

Observation 50ab1b23-a4ab-427b-86b2-c90363ac2245 · outbound

This paper cites P roof W riter: Generating implications, proofs, and abductive statements over natural language.

Reasoning Bias of Next Token Prediction Training P roof W riter: Generating implications, proofs, and abductive statements over natural language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.495140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.495140Z digest=sha256:842553f7b287ee9ba97631f84b497a448456fdd6265223d2c660d3ab29eb55d2

Observation 2d6c2844-7d4b-412d-b39f-f2f929cb2d8d · outbound

This paper cites Memorisation versus generalisation in pre-trained language models.

Reasoning Bias of Next Token Prediction Training Memorisation versus generalisation in pre-trained language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.500250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.500250Z digest=sha256:da67b86ecb7dcf660b38b9bf1700d391a9ec08ec3ba18166510661fa14181e1c

Observation 4d0b59b2-752d-417e-a8e7-4396418c8023 · outbound

This paper cites Implicit optimization bias of next-token prediction in linear models.

Reasoning Bias of Next Token Prediction Training Implicit optimization bias of next-token prediction in linear models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.547038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.505179Z digest=sha256:c0ff31edafa36cf4eaee06665fdeaf18a9756fad5b43069338583523fcb6eadc

Observation 9104245d-a585-445a-9e04-f9e73cdc3f99 · outbound

This paper cites An empirical study on robustness to spurious correlations using pre-trained language models.

Reasoning Bias of Next Token Prediction Training An empirical study on robustness to spurious correlations using pre-trained language models

Reference 45

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.713990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.509922Z digest=sha256:25680a0fb7ba66e670ccbbc56ad2bc5c2bb6fd8978b02dc47e50d3b3b874fa02

Observation 4d3139d1-8c38-47b7-a762-b88b682e82df · outbound

This paper cites L ogic A sker: Evaluating and improving the logical reasoning ability of large language models.

Reasoning Bias of Next Token Prediction Training L ogic A sker: Evaluating and improving the logical reasoning ability of large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.514920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.514920Z digest=sha256:737f0a095dd589d67c99d086ee797b61639b898f1ea89dcf3dd268bc24f5767d

Observation 9d3daf19-287c-4f20-aa3b-8b685bcc4b2b · outbound

This paper cites Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023.

Reasoning Bias of Next Token Prediction Training Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.529595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.519722Z digest=sha256:36bcdb23de2d35366601db5ef57b0ebeb5ebf0229fec0d6583e3dbcf0fce4c7f

Observation cc99b161-f150-445d-a9f1-c18593e4f879 · outbound

This paper cites The implicit and explicit regularization effects of dropout.

Reasoning Bias of Next Token Prediction Training The implicit and explicit regularization effects of dropout

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.513937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.524447Z digest=sha256:486dad4fcf41613f490e00346c08af2948574c1e67c166b9109207264aff6603

Observation 0448aee8-c29c-46bb-825d-985cba3b0ada · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Reasoning Bias of Next Token Prediction Training Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.529135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.529135Z digest=sha256:c4a41dbfe7cd449db42c84ceb56affdf644453cef99a55da186ab603755a0b60

Observation 7fcd35b3-2ada-4779-b3dc-6ee73103d0c0 · outbound

This paper cites On the noisy gradient descent that generalizes as sgd.

Reasoning Bias of Next Token Prediction Training On the noisy gradient descent that generalizes as sgd

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.496363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.534025Z digest=sha256:89f3bd80bd447d220ba1b5ca17ad8b374c8055318f3ffd3842accfe32fcd2a5a

Observation 21ebed3d-a2d9-431a-a33c-299eba61061f · outbound

This paper cites How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective.

Reasoning Bias of Next Token Prediction Training How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.480659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.539227Z digest=sha256:2ab17c87626307675144dd17061fbeea5236522afd7ea043eec6ce9b12e4713a

Observation e45537e8-832f-4ceb-9a8b-bfad031b41e3 · outbound

This paper cites UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost.

Reasoning Bias of Next Token Prediction Training UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.544075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.544075Z digest=sha256:66edb919111b58ef87eda61b602be092165024e14f4598bdec582acb1d12eaa2

Observation 87d8e884-4966-49fa-aae4-9ba4027f9c3c · outbound

This paper cites D., and Potts, C.

Reasoning Bias of Next Token Prediction Training D., and Potts, C

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.549000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.549000Z digest=sha256:8225461da4d31abb04a234ad6ac4a27d7d1eabd845801e6e636f8fbefd8235ef

Observation 93be4192-acd6-4e66-9346-1ba25f5a8a5f · outbound

This paper cites A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Reasoning Bias of Next Token Prediction Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.554031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.554031Z digest=sha256:9186a7c267a8a1ca169c04fba69417ccffb15f93636a877ee613e0a3c99633ae

Observation ef9ae870-53ef-4629-90b6-895e2c03ed87 · outbound

This paper cites URL http://www.yelp.com/ dataset_challenge.

Reasoning Bias of Next Token Prediction Training URL http://www.yelp.com/ dataset_challenge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.465135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.559095Z digest=sha256:30726ef2e33fdbec7b81805f645e1405cfe38399e0bd2a749f499de6cefa9fd0

Observation ffb8e00d-fa33-4783-b3d8-65f358a901f5 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Reasoning Bias of Next Token Prediction Training InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.563724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.563724Z digest=sha256:1813cabfd717a9f8f1dd2023f2cbe744f5a3b34ca09113f9dfb2e30da13b8130

Observation c837e3f2-6754-489a-a334-e4e3b937e06f · outbound

This paper cites Natural language reasoning, a survey.

Reasoning Bias of Next Token Prediction Training Natural language reasoning, a survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.568890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.568890Z digest=sha256:f5c9bf9828ec49349dca01d5077bf68b849315345c8d3ffbe6a015f29d2dd57a

Observation 7752d722-275c-47e2-a23a-c363ea5c1b53 · outbound

This paper cites DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks.

Reasoning Bias of Next Token Prediction Training DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.573794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.573794Z digest=sha256:92f4660e6bd0c39097b72821d8553218f0895b86b511094e7c16b00a71ab7856

Observation 3b913036-4a56-4b15-87c9-92fa484ac6e2 · outbound

This paper cites Dropdim: A regularization method for transformer networks.

Reasoning Bias of Next Token Prediction Training Dropdim: A regularization method for transformer networks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.579115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.579115Z digest=sha256:9428dccc3a3b9df241b403da770a670ce442442cce70d0ba1f9a8b39575b6374

Observation 74d55690-5c9c-4951-b325-91bfac772083 · outbound

This paper cites H., Meng, T., Chang, K.-W., and Van den Broeck, G.

Reasoning Bias of Next Token Prediction Training H., Meng, T., Chang, K.-W., and Van den Broeck, G

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.583817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.583817Z digest=sha256:8af6e88eb5c36d508e79e24e0db239835e7179f38c10505fba1b36e54c9a12a7

Observation 33ebdccc-6c5d-42b1-8b4f-0d681e7bd3f0 · outbound

This paper cites and Xu, Z.-Q.

Reasoning Bias of Next Token Prediction Training and Xu, Z.-Q

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.588944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.588944Z digest=sha256:4d32e76850788094e1d515d1e439d72de03c03ba57619305a8ab55891911ceee

Observation f61a50ba-e91a-4684-979c-b1dcc1ae298b · outbound

This paper cites Stochastic Modified Equations and Dynamics of Dropout Algorithm.

Reasoning Bias of Next Token Prediction Training Stochastic Modified Equations and Dynamics of Dropout Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.594201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.594201Z digest=sha256:ec43006501ff9b41da035b70274f77bb6779373a8b98ed35eba161a6fc664299

Observation 98a53005-a589-4721-9d72-b65c5d036bc9 · outbound

This paper cites Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing.

Reasoning Bias of Next Token Prediction Training Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.599362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.599362Z digest=sha256:b9ef22bb5de37882307198f1217264fb8b022662029a31cb55ac0df92cd2181e

Observation bc20519d-1dad-4b65-ae7c-c7d1f8a0e807 · outbound

This paper cites Anchor function: a type of benchmark functions for studying language models.

Reasoning Bias of Next Token Prediction Training Anchor function: a type of benchmark functions for studying language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.604752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.604752Z digest=sha256:bf7db0c94f6f88cae928bbb2feb5d890b35e80b1d1e8024b296fc65ab25aa40a

Observation ddbdd9be-60e5-4cbb-8e00-b098c371f2df · outbound

This paper cites Implicit geometry of next-token prediction: From language sparsity patterns to model representations.

Reasoning Bias of Next Token Prediction Training Implicit geometry of next-token prediction: From language sparsity patterns to model representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.440791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.609837Z digest=sha256:8eb03bd63d0ab0324411ff76aa9c9a77ec93f63ccb22acebaafe0cf3924d58b8

Observation d9675f5f-907b-4fcc-814f-96896e1c0ded · outbound

This paper cites Scheduled D rop H ead: A regularization method for transformer models.

Reasoning Bias of Next Token Prediction Training Scheduled D rop H ead: A regularization method for transformer models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.614551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.614551Z digest=sha256:13e0324260ef5cd6ff2053e22f2070e5f97877715026da869ac231f0ef6377eb

Observation 30af6bad-2d3a-4fcf-9b1f-dc57d4457767 · outbound

This paper cites The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects.

Reasoning Bias of Next Token Prediction Training The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.424735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.619470Z digest=sha256:b01d6a07a35b1ef0a468766bc6571e508765454cd04e51f4fee2d23504e348f8

Pith citing papers

Observation 291098ac-6303-4e3b-9c1c-f6473047bec8 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Reasoning Bias of Next Token Prediction Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:43.217179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:43.217179Z digest=sha256:04942dbe630839aa170104b13d305a490a1e88c3276e20b0d77c5e756b7f7d39