Pith. sign in

Paper Citation Record · LEDGER

Enabling Autoregressive Models to Fill In Masked Tokens

As of 9 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2502.06901.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06901 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:10:39.863142Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T02:05:18.834924Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T02:05:18.893265Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 45ceeb89-ee0b-437b-8672-b55d2db58788 · outbound

This paper cites write newline.

Enabling Autoregressive Models to Fill In Masked Tokens write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.679309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.679309Z digest=sha256:066f7733d832508490428715d15a0c67f47c63176a2fe9a794a64d35000cf112

Observation aa9b1412-ec3b-4e36-9a12-e63d306dff77 · outbound

This paper cites GPT-4 Technical Report.

Enabling Autoregressive Models to Fill In Masked Tokens GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.683696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.683696Z digest=sha256:3108c58330c1b64bfe4f9aeb18d2e896aa22b63786a7158d9ba5d6b0e644d5b6

Observation e573142e-5f56-4e34-b7c8-e45e4447d3d4 · outbound

This paper cites and Tsitsiklis, J.

Enabling Autoregressive Models to Fill In Masked Tokens and Tsitsiklis, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.366042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.687735Z digest=sha256:d123b3d76fb6ac8055d44ecf54f00532c9c56c6320fd5a5d89f9106aea9afc2c

Observation 01600c16-c573-41cb-a439-fcd9e478fa0d · outbound

This paper cites One billion word benchmark for measuring progress in statistical language modeling, 2014.

Enabling Autoregressive Models to Fill In Masked Tokens One billion word benchmark for measuring progress in statistical language modeling, 2014

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.355029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.691337Z digest=sha256:0055e3dcc82528eda3f181a9aab36c0bfc9115571f57b90f2618ee0c8ef6aa4b

Observation 58e772f8-d8c5-4d81-9a88-57d93e7104db · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Enabling Autoregressive Models to Fill In Masked Tokens Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.695329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.695329Z digest=sha256:6a2fe6d4d30e18063528c8b403fc3b14e6078ec13819132017241bc377f0fd83

Observation c4520c0a-6d43-4058-8f5a-7ae5d29b1778 · outbound

This paper cites B., Bierbaum, M., O'Keeffe, K.

Enabling Autoregressive Models to Fill In Masked Tokens B., Bierbaum, M., O'Keeffe, K

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.699400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.699400Z digest=sha256:7aed53d328bc41a503b6066fecad7f55e264695a2aab0397be3296fa629dc7e1

Observation aa22d8de-870e-4f6b-af3e-971308a88e9a · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Enabling Autoregressive Models to Fill In Masked Tokens BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.703116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.703116Z digest=sha256:5de58ad84d19f7ddc5d364de4cc59e21177d780986fc3670bbab62a88b642a5d

Observation e54ac4dd-2c52-418e-bce5-dc530a7ef369 · outbound

This paper cites Enabling Language Models to Fill in the Blanks.

Enabling Autoregressive Models to Fill In Masked Tokens Enabling Language Models to Fill in the Blanks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.707483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.707483Z digest=sha256:c039d43e3f9985067c9c48da80ecb38ea8fa8c6ca717755f8445409b72902488

Observation 9a255a77-acb0-4288-bba8-2ddfc9be5f11 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

Enabling Autoregressive Models to Fill In Masked Tokens GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.711429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.711429Z digest=sha256:68846f62b041dfe722b4dd0808eca160c6fd1a7ee416a603eaab5ebf9e41bd55

Observation 4242e82c-bb5e-4a14-ace2-bc95bdb7a307 · outbound

This paper cites InCoder: A Generative Model for Code Infilling and Synthesis.

Enabling Autoregressive Models to Fill In Masked Tokens InCoder: A Generative Model for Code Infilling and Synthesis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.715104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.715104Z digest=sha256:30794115cbfa061df9dc90daa6a5cf7ab170097f8975bb17a48044f5fda8860f

Observation 24f7c9a0-1632-4011-acfd-8a875586ff9b · outbound

This paper cites Scaling Diffusion Language Models via Adaptation from Autoregressive Models.

Enabling Autoregressive Models to Fill In Masked Tokens Scaling Diffusion Language Models via Adaptation from Autoregressive Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.719321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.719321Z digest=sha256:0c5770a32a4d0d571cfab9fdd001af510878120dd469fbf53f3665318b262f8e

Observation dbbcb1ea-bf3a-4385-8f2c-db9ba5aa9370 · outbound

This paper cites The Llama 3 Herd of Models.

Enabling Autoregressive Models to Fill In Masked Tokens The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.723058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.723058Z digest=sha256:571bae8058c62c9745b7635cc773668e68d5705a3a62f5e5bd38c3d502e939b5

Observation 4eb12ada-59d4-4657-acfb-ffc2535bcd7f · outbound

This paper cites OLMo: Accelerating the Science of Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens OLMo: Accelerating the Science of Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.726249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.726249Z digest=sha256:640375e9ee6e8f9c6717602a200103826978d86408a159f8d73c2e08d98e3ca2

Observation e193fb6a-9c63-48e5-9e23-47c67a05afbe · outbound

This paper cites Likelihood-Based Diffusion Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens Likelihood-Based Diffusion Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.729397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.729397Z digest=sha256:3aaae1ecd4aab61850a8ef331d0c81240761bed9653e084e95905197f91dc46b

Observation 48ee7a31-8f0d-411e-9470-6010de50c662 · outbound

This paper cites an unresolved cited work.

Enabling Autoregressive Models to Fill In Masked Tokens Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.732855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.732855Z digest=sha256:27fd32a4f494e0e0e47aff50f3a440693c73d54c2da4853ffe2979367dd0e208

Observation 968eef7d-538b-40c6-9618-97a694d26076 · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Enabling Autoregressive Models to Fill In Masked Tokens Denoising Diffusion Probabilistic Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.736815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.736815Z digest=sha256:601922bced910df5cfc3be088f4d34667f0a983a1c3888a61e05b52fe1d56c7a

Observation d1ebc4c2-b695-4d4a-a0a3-88643b99ee76 · outbound

This paper cites The Curious Case of Neural Text Degeneration.

Enabling Autoregressive Models to Fill In Masked Tokens The Curious Case of Neural Text Degeneration

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.740563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.740563Z digest=sha256:9324ce4dc139cabfe3beed69274217e495904017658db2510df239a633e3d02f

Observation 7e2421f2-f268-4ed5-a450-08f6fb4d9dc3 · outbound

This paper cites Autoregressive Diffusion Models.

Enabling Autoregressive Models to Fill In Masked Tokens Autoregressive Diffusion Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.744140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.744140Z digest=sha256:97a9b3887912ec30988c52ef41f3de11687aa89b27235cbf83dacdefcf21ef68

Observation e62cf978-a81a-4b93-abe4-ade911d4ec97 · outbound

This paper cites Variational Diffusion Models.

Enabling Autoregressive Models to Fill In Masked Tokens Variational Diffusion Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.748997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.748997Z digest=sha256:153c4384ba9ff3a46a7d5fcc93f9e6cccbd6f26e2ecf9581a09f13b05f717eda

Observation fb27396a-9ecc-4e2f-a49b-3978f4bd7d2d · outbound

This paper cites H., Gonzalez, J., Zhang, H., and Stoica, I.

Enabling Autoregressive Models to Fill In Masked Tokens H., Gonzalez, J., Zhang, H., and Stoica, I

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.752819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.752819Z digest=sha256:815ca1e0fd8b430285411477f0e1ab50c349265a973985437b52193f80682fc2

Observation b886ea01-81ee-45fd-a766-ccd00f46ae64 · outbound

This paper cites Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities.

Enabling Autoregressive Models to Fill In Masked Tokens Coauthor: Designing a human-ai collaborative writing dataset for exploring language model capabilities

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.323480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.756278Z digest=sha256:f1d3d98361bd0dbfdfcaacf22824376f93a6acb1eb45813f66b83b0234334872

Observation 92705d52-cb0f-430d-9ecf-6c84ddec48a3 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

Enabling Autoregressive Models to Fill In Masked Tokens BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.760569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.760569Z digest=sha256:ff15ede9c2011ff431537d686c35403b680c8bda3de696758c92404f76d8fffa

Observation 6c066e2a-a343-4593-8d26-203564067ceb · outbound

This paper cites DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts.

Enabling Autoregressive Models to Fill In Masked Tokens DExperts: Decoding-Time Controlled Text Generation with Experts and Anti-Experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.764528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.764528Z digest=sha256:747e94179dd129380ea7c267c3db152f576e7692b8ebb44e93cbe1bc6ec5fce5

Observation 5de9f587-e63d-415a-8319-08328890cb91 · outbound

This paper cites Discrete Copula Diffusion.

Enabling Autoregressive Models to Fill In Masked Tokens Discrete Copula Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.767786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.767786Z digest=sha256:9a834f2c1ff81ff69813b2877b8b1ae13567da2965ab15f5c634b00115c9e413

Observation d86f4111-c2ca-4400-9506-c771462a25e6 · outbound

This paper cites Multi-task learning based pre-trained language model for code completion.

Enabling Autoregressive Models to Fill In Masked Tokens Multi-task learning based pre-trained language model for code completion

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.312183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.771189Z digest=sha256:a1fde5d6e6ebed1455883eb369fe83b6128a8487e266e4e12ec2c5cd2886c8d8

Observation c40bb9f5-b970-4f5e-9c2a-72be2c6c62d1 · outbound

This paper cites Think While You Generate: Discrete Diffusion with Planned Denoising.

Enabling Autoregressive Models to Fill In Masked Tokens Think While You Generate: Discrete Diffusion with Planned Denoising

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.774186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.774186Z digest=sha256:9435ba9ef1c1e3ae4dfc2634248b517956e29d6ac098cff9d6389954ca2680d4

Observation 5f49b491-0638-4fe5-b7aa-7f6189e74055 · outbound

This paper cites Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time.

Enabling Autoregressive Models to Fill In Masked Tokens Scissorhands: Exploiting the persistence of importance hypothesis for llm kv cache compression at test time

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.300241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.778773Z digest=sha256:823b5382e3a9cfc55d218fcf38b65bbcc728a185ce6468f10200050b87359874

Observation 06bb53cc-98f8-4a9d-aa0d-3a1ef204a551 · outbound

This paper cites Discrete diffusion language modeling by estimating the ratios of the data distribution.

Enabling Autoregressive Models to Fill In Masked Tokens Discrete diffusion language modeling by estimating the ratios of the data distribution

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.288372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.782296Z digest=sha256:b6b320a3eddeb994fd9b8ad3391308700e0fc8db64fecb5a32cf2e1380895f55

Observation 5e549bb7-c074-4fa7-9bc8-89c595d637e7 · outbound

This paper cites an unresolved cited work.

Enabling Autoregressive Models to Fill In Masked Tokens Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T17:10:40.275356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.786049Z digest=sha256:6f638dfd34aad59ac404a7b3400927b7f50a134f50ab00fcd66259a715fd8437

Observation 9143eb05-2bda-4d7d-967a-f69500c7f729 · outbound

This paper cites Pointer sentinel mixture models, 2016.

Enabling Autoregressive Models to Fill In Masked Tokens Pointer sentinel mixture models, 2016

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.789928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.789928Z digest=sha256:626e16b75822966a4ed7c5215fc6136f7bb14f3e9e4e7e6f0d2a2fd6987ca8ac

Observation 0d302c15-0ae8-4383-81bc-c27efeb2e2ff · outbound

This paper cites Meet in the Middle: A New Pre-training Paradigm.

Enabling Autoregressive Models to Fill In Masked Tokens Meet in the Middle: A New Pre-training Paradigm

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.794278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.794278Z digest=sha256:743d46fba9a89ed2b9324911e59effa1b8eaf934b3ab8ce613781ce00985fd51

Observation 6eb88fef-74fb-480a-af46-1c9c9499e48d · outbound

This paper cites Scaling up Masked Diffusion Models on Text.

Enabling Autoregressive Models to Fill In Masked Tokens Scaling up Masked Diffusion Models on Text

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.797937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.797937Z digest=sha256:e06b54e45e9821f4397930c08f79d7f895a97217542cd3ba665322d47efacdef

Observation acc9f699-998e-42db-9012-1388bd6c97a9 · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

Enabling Autoregressive Models to Fill In Masked Tokens Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.801915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.801915Z digest=sha256:6a8233910ede2f8128d129676e979b597414848256a04e11097faf429a20069d

Observation 28bc4620-320d-40f1-b56c-8b60d4f1b549 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

Enabling Autoregressive Models to Fill In Masked Tokens The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.805621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.805621Z digest=sha256:0f45570e10d25fa41ef79400cbc3da500d18f3749c814858e2133f0306bec37c

Observation 32c2ace4-0804-4bff-9f5a-8690dabc4b0a · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

Enabling Autoregressive Models to Fill In Masked Tokens The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.809561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.809561Z digest=sha256:29827e5b348874533db278f8ee7c8e1ee708bcba39213d2d080e53372d0b12d3

Observation 38ef1f46-d3f0-45da-aa88-1124671e7d22 · outbound

This paper cites Language models are unsupervised multitask learners.

Enabling Autoregressive Models to Fill In Masked Tokens Language models are unsupervised multitask learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.813372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.813372Z digest=sha256:166661ddc184609eddf6db72f30c2673bd163723a1ddf745dd01b60579c1c3d4

Observation 025c07d4-cfa7-459f-8ff5-a347b04b23df · outbound

This paper cites Simple and Effective Masked Diffusion Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens Simple and Effective Masked Diffusion Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.816372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.816372Z digest=sha256:e1dcbf46eec6a0a47066712e0bed2d18d96c1bad2084e87dcd1cff856830a642

Observation 1026cd01-100b-4f3a-8d0a-91fd78437d97 · outbound

This paper cites Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization.

Enabling Autoregressive Models to Fill In Masked Tokens Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.819680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.819680Z digest=sha256:2e92f238637676ee738f4ab5d5c0181b5a92b3c9d4c05e05971e0084f4a0bf22

Observation 3bb74a02-3c28-4a66-999c-a45a4e69665b · outbound

This paper cites BERTs are Generative In-Context Learners.

Enabling Autoregressive Models to Fill In Masked Tokens BERTs are Generative In-Context Learners

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.822853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.822853Z digest=sha256:cd2e0492f1be918ed5f733ebebcc56b44eae5333954c2224e69669be88588024

Observation 81ac30c0-0010-49de-a9b3-87e796c9129e · outbound

This paper cites Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition.

Enabling Autoregressive Models to Fill In Masked Tokens Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.825969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.825969Z digest=sha256:46e303077273a4265537f34777cb2a8c8d5cafbc6c9d890877833f1c2edb57fb

Observation da39abf7-16ed-4c1d-b224-23d37adaedd7 · outbound

This paper cites FiLM: Fill-in Language Models for Any-Order Generation.

Enabling Autoregressive Models to Fill In Masked Tokens FiLM: Fill-in Language Models for Any-Order Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.829805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.829805Z digest=sha256:6a2efb3368e4be989ef5a1f9a215996c0ce3b914c7914a44b2ccd4586d1b9d74

Observation e5b9d0ce-9604-47de-8fee-e56bc88f3b1e · outbound

This paper cites Long horizon temperature scaling.

Enabling Autoregressive Models to Fill In Masked Tokens Long horizon temperature scaling

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:10:40.250448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.833475Z digest=sha256:c339c060ed881e626c54dedc5d0aeecf68154dd343c2d2748a8177d53ae480c1

Observation 96e8022d-ec5e-44db-9812-ba330e4e52ba · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens LLaMA: Open and Efficient Foundation Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.837397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.837397Z digest=sha256:d1fdd4d22bacea5d58f07454b04c26c9aaca116112fa2abf79b678768c7ee8ee

Observation 59df9685-7ca3-4afe-8594-691462055097 · outbound

This paper cites Attention is all you need.

Enabling Autoregressive Models to Fill In Masked Tokens Attention is all you need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.840865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.840865Z digest=sha256:a59f6b8f14d0f228276286e41fc98da35082bb9a7fa68b42315ad4b45a72f1e0

Observation bd2e9aa9-c568-4098-a642-394e4aa87d57 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

Enabling Autoregressive Models to Fill In Masked Tokens Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.844222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.844222Z digest=sha256:b20a19a20f15e0af4d58f6d3f319c7305fbd2b5ae83ad51bec24582708e36716

Observation 449781f1-aa04-4d85-9d85-8fefacf44433 · outbound

This paper cites FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability.

Enabling Autoregressive Models to Fill In Masked Tokens FOFO: A Benchmark to Evaluate LLMs' Format-Following Capability

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.848115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.848115Z digest=sha256:2629597ef5686bfaa680a7b42c96cb190ab95790672ea4fa510a67cbaf7479d7

Observation 07fc3082-f54e-4699-8617-996902045114 · outbound

This paper cites AntLM: Bridging Causal and Masked Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens AntLM: Bridging Causal and Masked Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-08-08T17:10:39.931496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T17:10:39.851945Z digest=sha256:f13ef57bc3d752129e169442f86acea39eba82a0bb584873fe3ecef1dd032161

Observation 90600637-0f5f-4e81-90aa-9ccd8a3b7a98 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Enabling Autoregressive Models to Fill In Masked Tokens Character-level Convolutional Networks for Text Classification

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.855961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.855961Z digest=sha256:fd9f2ffab6266898d36c6de382bed2ce1f2800e2d0f41873231f9b14c5e5809a

Observation 56c57e3a-696d-4189-b0b7-dea6c0b062c8 · outbound

This paper cites Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens Prepacking: A Simple Method for Fast Prefilling and Increased Throughput in Large Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.859709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.859709Z digest=sha256:0f91b0833d28cd62eb4ecc2023d500e545baaf1f79f046b5bd4c4a44be4836e9

Observation 4658089a-70a2-4a5b-a299-6077e0b480db · outbound

This paper cites A Survey of Large Language Models.

Enabling Autoregressive Models to Fill In Masked Tokens A Survey of Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T17:10:39.863142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T17:10:39.863142Z digest=sha256:a2b1a09b9f641ce3b5796e6b09d08610a1590b556d602fa2a58f0a44f8a1804b

Pith citing papers

Observation a569b31d-131c-4fda-9330-69ec8d02b2fa · inbound

Mercury: Ultra-Fast Language Models Based on Diffusion cites this paper.

Mercury: Ultra-Fast Language Models Based on Diffusion Enabling Autoregressive Models to Fill In Masked Tokens

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T02:05:18.895176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T02:05:18.834924Z digest=sha256:c7b860397f77019cc769ed5575cd895a4ccffbbf22e0c719174e368fc2189fa9