Pith. sign in

Paper Citation Record · LEDGER

On multi-token prediction for efficient LLM inference

As of 8 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2502.09419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09419 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:38:35.138928Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e1c89d7-4123-4a29-b3d4-e9397a08b79a · outbound

This paper cites Llama 3 model card.

On multi-token prediction for efficient LLM inference Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.611805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.611805Z digest=sha256:924d5a43cffeb4f9cdf54ebf20bbaee986a11004578041d24856f34f0d3f73d0

Observation 63302489-244a-41a6-ad7d-cbf2bcfede47 · outbound

This paper cites Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition.

On multi-token prediction for efficient LLM inference Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.615153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.615153Z digest=sha256:73c4ef6ec303511ddac987b26506b10024206492689e9b4f13439ae12fecfad0

Observation cd1dd56f-b51d-4fc3-8f37-6ac15e6d8f99 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

On multi-token prediction for efficient LLM inference Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.617991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.617991Z digest=sha256:c1898ff8128d0acd5f8e705105670b85c7142a66d8e6679503d63c00b816e8b2

Observation e64287f3-1e88-4e5b-b48b-0f970b9dfa8c · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

On multi-token prediction for efficient LLM inference Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.620399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.620399Z digest=sha256:a799ad955dc0ce6d193b3695df28a82d8df02ca02670482ba63039e7fb6379c8

Observation abc25476-d4c4-4961-afff-c60c89533ed5 · outbound

This paper cites Overview of the IWSLT 2017 evaluation campaign.

On multi-token prediction for efficient LLM inference Overview of the IWSLT 2017 evaluation campaign

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:38:35.309185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:38:34.623866Z digest=sha256:2b9d399d98b9b7ee15fde2a603edfe7e3bdec94c75b22ca127d28f0963ca3d1f

Observation d54e0028-cae2-44b5-8d96-da2099414230 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

On multi-token prediction for efficient LLM inference Better & Faster Large Language Models via Multi-token Prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.626760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.626760Z digest=sha256:c150acf8deafbacbd10c406bee4f20a9e03f94707e88d6133db125b943dc1770

Observation 6a4f7d3e-5c02-4737-8617-3a3bfe9b68ae · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

On multi-token prediction for efficient LLM inference LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.707368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.707368Z digest=sha256:b1de29d772012b70afdd66d2bd5aebd8d53bee79e360152bcf648e8343057a4b

Observation 4eee505b-e77f-4694-be24-ab55ae4e2279 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

On multi-token prediction for efficient LLM inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.760608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.760608Z digest=sha256:51965f0a8101cde4dc081b87f00f4a04495e649f94b405b26073c0a9333d98af

Observation dbb0d9f6-7754-44b0-88da-a21342b2432c · outbound

This paper cites write newline.

On multi-token prediction for efficient LLM inference write newline

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.861556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.861556Z digest=sha256:5cc39e0a30ae3af1a826935cc60750ac1b73604be4c9dfbf1aa012ac6c869bcf

Observation 1e0ca3e2-175a-48f9-8738-c2b90a54b7d9 · outbound

This paper cites @esa (Ref.

On multi-token prediction for efficient LLM inference @esa (Ref

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.929669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.929669Z digest=sha256:02d85cbff03d44a59c760530c6ccfc7586b05b4d9653ceeffe0fc801caebad9b

Observation 17263872-a3ca-4f77-882a-ed1c88d1b384 · outbound

This paper cites an unresolved cited work.

On multi-token prediction for efficient LLM inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:35.011306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:35.011306Z digest=sha256:1b5c4dfed45a002dfc3105e4d2547ff4b2fd7bd27ef6574beba185f161c309d3

Observation 87a04758-82ba-4bb8-84f4-32199623e755 · outbound

This paper cites an unresolved cited work.

On multi-token prediction for efficient LLM inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:35.138928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:35.138928Z digest=sha256:92143b0aaf7a81a1f327c9d220e3189eceafe87df29efda2623094a508f0782a

Pith citing papers

No inbound Pith citation observations are available.