Pith. sign in

Paper Citation Record · LEDGER

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree

As of 13 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2412.12639.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.12639 v3

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T13:56:39.808556Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-12T01:08:14.247342Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:26:24.739905Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved45
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 088995a7-78ae-4a33-96d8-086b3ba695a1 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.544245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.544245Z digest=sha256:6bf4fe24322b339d2be47e2cbdba01eea822bf1bf13d399c6a2c7fc43d860c6e

Observation 1095de3a-02e5-48b5-8e39-cdeb02d13c4d · outbound

This paper cites write newline.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.549623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.549623Z digest=sha256:e15db51a7ff3bcc4c6b7712cec66391aceba7c6b2b67abc742e793f676ae3701

Observation cd21b300-3a5a-4f28-bccd-710ef8584d9c · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.560090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.555810Z digest=sha256:01f621757bc5047a835bdc6ab11e7dc0906b72f21d86cfaeed7db5eeec31bb83

Observation 1f4542ba-b36b-4454-b281-de72137be92f · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.561574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.561574Z digest=sha256:f6feddc9411a639b787f9cf870095e1efc3342a502156008e2e88985d346f3ef

Observation c8fa522f-0984-42ad-99c2-8f573c815683 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating Large Language Model Decoding with Speculative Sampling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.568070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.568070Z digest=sha256:510555064d3d58e370ab3bd0eb4710ab355de9623b81b58ddb9bb570266a913a

Observation b102a94c-ac3e-486c-b195-872c3f3f1a09 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.573168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.573168Z digest=sha256:7a1b56a37b6aab4dab114056464cd769e5e4eef4b01f8874ce710054e6d7dcf0

Observation 85d3b22d-8c13-4fad-9bc1-0a8201207dca · outbound

This paper cites Cascade Speculative Drafting for Even Faster LLM Inference.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Cascade Speculative Drafting for Even Faster LLM Inference

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.580987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.580987Z digest=sha256:1e94681741022bd44ebe5dfa82609bc0605bc16532d18c106607a28d3c3cab74

Observation 49babc3b-b03a-4f88-9090-b3237ccf8188 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Training Verifiers to Solve Math Word Problems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.586758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.586758Z digest=sha256:d5e26e56fe203dd32c35b00ef99bf1c335a0f021a9e8a8fafb4acbf68cdfcded

Observation 330b6d1d-7df0-4e21-8124-61eeafdf9cd6 · outbound

This paper cites F.; Tao, D.; and Tu, Z.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree F.; Tao, D.; and Tu, Z

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:56:40.547022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.592117Z digest=sha256:63f24c4bba28e57c7892bfe3d3fe59aabb369e9f6a1e862b24585063885cf487

Observation dfc0bfd4-2231-4a0d-b1f1-898a3cd10303 · outbound

This paper cites GLM: General Language Model Pretraining with Autoregressive Blank Infilling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree GLM: General Language Model Pretraining with Autoregressive Blank Infilling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.596093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.596093Z digest=sha256:2785fdaa4ed016b275236629472d79d13207e3d103df9a30557c32b62149cb0d

Observation 2ada1f50-1235-4c8d-a9bc-f04c24a8fe48 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.534753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.600341Z digest=sha256:008694ec18b52063be8bb31c0dba57f86c2e408e9de348855f7fc2dc2efa1380

Observation 6f84257a-97a4-414e-9fef-66e472a28cbf · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.519615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.603874Z digest=sha256:f4b78d38bdcd0c41450a6f10cb12b7a3268c2b0370742598f589f2d050463efe

Observation 175d18a3-778f-45f0-a3f4-316ff51a0d06 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Better & Faster Large Language Models via Multi-token Prediction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.607525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.607525Z digest=sha256:fdc0d5f0cd6fdb8d5b81c20f6affc5d91531bb77ce6774ab93f43b5746464b72

Observation af66018e-a323-43ff-a041-e8aa6ef34276 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.612401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.612401Z digest=sha256:527196431fcf4c48b7407027bf9afe7d2440abb85551295e3f6f611d67a581ee

Observation 8ab821b5-56ec-4d84-af51-00a24bbfdb33 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.493299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.616664Z digest=sha256:d91cb26d6e2e8ec908fe56d5b9bac0d59923f7e5abe39610217009fd28de1ab3

Observation 8fa911fd-5676-4554-b832-7d0412bbc4be · outbound

This paper cites Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fine-Tuning by Curriculum Learning for Non-Autoregressive Neural Machine Translation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:40.085206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.621807Z digest=sha256:28a5c8f80182f469399cb91fe8979b66ca25874c112ad96ccdcad5c4abd84fef

Observation a288e5bd-b530-420d-939f-c51e55f907c5 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.477145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.627068Z digest=sha256:0ea1470528ceb0acbcccf548c54c31334bf18e83a1853757c65aca12d71099e2

Observation 4482b961-7191-436d-a1f9-4a51fd5f0c59 · outbound

This paper cites Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Self-Distillation Mixup Training for Non-autoregressive Neural Machine Translation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T13:56:40.061510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.631862Z digest=sha256:7f884f02f006ce2ee53fcbb886ab54e33ed10690e9a8974dd149d95f30f10137

Observation e903e635-f4ba-40a0-a7b8-19bfd2aff34f · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Distilling the Knowledge in a Neural Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.638681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.638681Z digest=sha256:ae0e7a490905e900faa754b1d66fa47287f81e821accac991f1672c6bc724bf6

Observation 47ab6896-f1db-4b66-b3b6-ac9762015293 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.642790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.642790Z digest=sha256:71edc7be7a16bbb0af87bfc26bacae7f0a3e7ef49191ea7a82d3d1b35ffb54d8

Observation 176e4f2f-8909-4647-a87e-36107980e473 · outbound

This paper cites W.; Gholami, A.; and Keutzer, K.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree W.; Gholami, A.; and Keutzer, K

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T13:56:40.446329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.647897Z digest=sha256:123eb1334d2a430f96f06969f8c514b4b43a9d507711f4b034c289871e661a0b

Observation ba1f2e9a-4969-45f3-9bfd-3d67295ac2cc · outbound

This paper cites Fast Inference from Transformers via Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fast Inference from Transformers via Speculative Decoding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.653053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.653053Z digest=sha256:0de9ba3dcb78e987e5320d8f5798baf815502b1de844e87f08586f9cb281d068

Observation 018023f6-a511-40a7-abf7-378e706924c6 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.663532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.663532Z digest=sha256:1b378dfe93d737e03ae0e1a602fc12fdbe7b8b0843cce34352b4893a1cc926b1

Observation fd32a44b-b5b5-45b3-80be-770b28e89738 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.426314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.670302Z digest=sha256:28ae8f1b059204e66f4d0c67955fc9f66c782e95520ff36c2553b93ef8852e6a

Observation 9c9e0826-8656-4623-89a3-6b8b1145ee52 · outbound

This paper cites PaSS: Parallel Speculative Sampling.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree PaSS: Parallel Speculative Sampling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.674571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.674571Z digest=sha256:b3d0f3c829045a2828ee3ca5b64777bd3b90139b8f5ba7d6f545883bf3847f56

Observation dc02a5c5-cb1a-44f5-b382-3ce5f730c8e1 · outbound

This paper cites In-context Learning and Induction Heads.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree In-context Learning and Induction Heads

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.679327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.679327Z digest=sha256:9600c4415f7de66c1050b85c6ad6eb062da13d1dfa6167e3a118738d2064131a

Observation 85628b68-1705-4171-bacb-ad0d145d294f · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.408816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.688977Z digest=sha256:782af1f452321b5e510bc3348c01e88642081c6c3a8fbdac3c95c918e8dcab8a

Observation 9885abe7-7969-4c1a-8386-d0338e15da0a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.388929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.693837Z digest=sha256:d007052bffdf7ccd260bfd2c4992969e68960744dd9a26896fe6d260ae4e895e

Observation d0236f5d-c9bd-4bdf-8d7c-04b16e9ac3d0 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.369184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.699417Z digest=sha256:09340cdfabbb1d9e232b47d332d24627643394ad8024e2d2375bbe40fd69bbc6

Observation 7ff372d0-6ce5-423d-9cf6-8092dfad27ca · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.348487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.707155Z digest=sha256:14d91023a0d9ea8fb40cca09e6f83216b514c83782a7e4a0366e3b22a1e6871a

Observation dd547bb5-cfd0-4fcf-9315-2da5ce487123 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.329362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.713796Z digest=sha256:d28ba9f3fb511f5ffbb3566b84d272c54b3a639ce5d584e8163b9c05a1766af9

Observation 975475f5-45a8-4cae-b9be-ff813764f66e · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating LLM Inference with Staged Speculative Decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.717832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.717832Z digest=sha256:881fb637d611add3baeb05cc1522b4c9d77f3517eb502d27a0a1323a4e5bc0b2

Observation 209fa9cc-c15f-45fa-a7bc-36baa9148d3a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.313913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.723725Z digest=sha256:b37478a72f4a192f7b6a300cce410fb8fc5a7f5bec97a66af98a7246b9371d54

Observation 62d5b023-e51c-4dfe-a105-c319e0e4f94e · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.298005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.729611Z digest=sha256:5fd78c15fc7b7665b110881cb9e791ce993fd4ae64f8933472431460586c5b1b

Observation 15dee101-6225-40c4-b916-5ceafb872b5e · outbound

This paper cites WaveNet: A Generative Model for Raw Audio.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree WaveNet: A Generative Model for Raw Audio

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.734235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.734235Z digest=sha256:d3c7abea46b3ba2844d38379323b8b694948c109ef875e151fcd77bf94a307d4

Observation 329f5adc-5855-4742-a60b-bdb427d69611 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.280552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.739625Z digest=sha256:ea274791fefdf883554b7672742ac2bf3b0df87402e3dc35843eaa83f9bc2f14

Observation 162cc3db-c0d0-45ed-b514-07c4426e776d · outbound

This paper cites Efficient Large Language Models: A Survey.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Efficient Large Language Models: A Survey

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.744899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.744899Z digest=sha256:7560ce59d6bf50d0487cc8f3e70dcd0bd6782f305a0645e8193ef483c8f69658

Observation 9f89cce7-c135-44d4-aa75-2e5694ae6d62 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.749541Z digest=sha256:34f2139e171ec2ec4748e540624fd698a77341786889bda129cb0061e20fe316

Observation 560487e9-aabd-4d9c-9d73-6e1b71cf50c8 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.242349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.756167Z digest=sha256:130251c727bf859ed2844189670cea4ace351888336e88cc1cf39c32a0833e9b

Observation 399bc475-9fcd-49fc-b503-f18391603f7e · outbound

This paper cites Accelerating Production LLMs with Combined Token/Embedding Speculators.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Accelerating Production LLMs with Combined Token/Embedding Speculators

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.760297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.760297Z digest=sha256:cf6fdab6960399f8cf4433ef6f0975ecd9cd1fc9addb748d60adb147070308f4

Observation 9e19fcaf-326a-4171-89ea-ff63a7fc60b1 · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.217318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.765552Z digest=sha256:67efc856b1e9262eb6e6c6b729033191a8fe1bb4055e7a5934529411b2b76f0b

Observation 6a36a5c1-318b-4044-b772-3322b34f97d5 · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.772067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.772067Z digest=sha256:f206e42eb7242c77bd94f10162c80c2b4e014647668a321467154974ea726980

Observation 355893e3-66d0-4fdb-9c4b-b7356dafd000 · outbound

This paper cites A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree A Survey on Non-Autoregressive Generation for Neural Machine Translation and Beyond

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.777354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.777354Z digest=sha256:737da03815758827c167aa01517b2fd538064a60c6e529d8f2ffa157591809a4

Observation eecc80ce-032a-4d7f-a8c0-6b2e2fff821b · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-11T13:56:40.202894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-11T13:56:39.782758Z digest=sha256:31787bb5a0365d55f4f273f37429aacdab3869754b56f3a8c5c79f3004195ed6

Observation 368bf258-9b1c-4e01-a035-843ea2b7ba09 · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.788351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.788351Z digest=sha256:d041b86d24529643aa33af274d0636166f0950bb0edd2c8e22a78a97fb2d311c

Observation 871eb455-5dcf-4a0b-8ac6-99956511008e · outbound

This paper cites Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Lookahead: An Inference Acceleration Framework for Large Language Model with Lossless Generation Accuracy

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.793794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.793794Z digest=sha256:3d61302ccd06ef8516501246ebb3ea2ab6af0cfdb370c38b9be7bb62e2928d32

Observation 1320dda8-e617-4b36-adfa-b0a7a0f8b86a · outbound

This paper cites an unresolved cited work.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.798834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.798834Z digest=sha256:6ed8a1f8d975efcfca4a812f6b1c772c23b08075ba35872f53a81faa8984db5a

Observation b884096a-6e81-40bd-ac00-1248f41f1541 · outbound

This paper cites A Survey on Model Compression for Large Language Models.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree A Survey on Model Compression for Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.803614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.803614Z digest=sha256:14b6b0b78e139b8651e8bd5417deaeea3e47fb9db5c94431b98d8904fa223d20

Observation 6908f2f6-4ebb-4d34-8076-e355979e08aa · outbound

This paper cites Fast Decoding in Sequence Models using Discrete Latent Variables.

Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree Fast Decoding in Sequence Models using Discrete Latent Variables

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T13:56:39.808556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T13:56:39.808556Z digest=sha256:439a7c68c26574bed6a628211a1c652c6f22b2de92d948be002db0bdfb27100a

Pith citing papers

Observation 4e8c7bb5-17af-47b6-a400-6099c7817c98 · inbound

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding cites this paper.

PARD-2: Target-Aligned Parallel Draft Model for Dual-Mode Speculative Decoding Falcon: Faster and Parallel Inference of Large Language Models through Enhanced Semi-Autoregressive Drafting and Custom-Designed Decoding Tree

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:26:24.741992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T01:08:14.247342Z digest=sha256:7d87e2e02792332d8e0cefcfe808b55c93202cea99475f3adb2b1304b3eb1cd1