Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:10.198594Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2505.23219.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:10.198594Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
41 of 41 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 24ae3333-f6f1-45e2-97c8-e76e0b08e213 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Efficient memory management for large language model serving with pagedattention,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91206771-a7bc-4a05-a2bb-0b33cdef305d · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Specinfer: Accelerating large language model serving with tree-based speculative inference and ver- ification,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9fb6f691-5aac-4615-bf7b-817a44725dce · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94627f43-b782-4c43-bcd0-91244600d07d · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Break the sequential dependency of LLM inference using lookahead decoding,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a371d8a0-6efc-4a5c-9ca1-5897462139c6 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient CPU-GPU co-execution for deep learning inference on mobile devices,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e581d66-b47b-4ca8-b845-521ed67b90f7 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0c9e5fbb-7dfa-4666-9e70-05434b1d00ee · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism High-throughput cnn inference on embedded arm big. little multicore processors,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32fc7452-b213-4759-93c1-5844b13f9b00 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Dopia: online parallelism management for integrated cpu/gpu archi- tectures,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2baf9dd-8904-47ae-9ed1-bba66802decb · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple m4,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a9dad62b-73f1-495f-8299-3ae6003b911f · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism vllm github,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 58b9bea6-9407-493a-97a4-732d593c3e7f · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cefeb681-f9e5-4b0f-bc37-fde6ef58283c · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7432d218-f45f-4fdd-8522-98a42ebbed96 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Intel core ultra processor family,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f3fc7d8d-41d0-46f8-9d9f-ff0a5f54c87e · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Meet jetson, the platform for ai at the edge,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f5589f97-86ef-487d-8f49-38482ae3dc32 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication effi- cient distributed machine learning with the parameter server,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f016866a-0884-484b-9494-5f9d5be016f1 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5a4f5d34-2367-49d8-981a-1361d9009a5b · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient cpu-gpu co-execution for deep learning inference on mobile devices
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9c66d894-9c9a-4eb7-98b3-69b0b4c28a22 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Attention is all you need,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5bb6186d-b056-45b1-953a-ddd20e647f17 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6096911-0308-49be-a5c8-4507020568d7 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism EAGLE: speculative sampling requires rethinking feature uncertainty,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f73d127-3366-4a33-8495-f3a5a873b60d · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple a17,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9580a79-5031-479f-be40-09730113aa80 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Amd reveals next-gen desktop processors for extreme pc gam- ing and creator performance,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0a0d0a69-c945-4d6d-b371-6b32b9cb5ad6 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Qualcomm snapdragon,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0ee5fff8-57fe-4f4d-8b57-00257fc7c042 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ring Attention with Blockwise Transformers for Near-Infinite Context
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29985fa0-3612-487d-8f4e-3f1164b86d7e · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Flashattention: Fast and memory-efficient exact attention with io-awareness,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9b6bcb50-3f38-4bdc-9901-6ba9cc0e51d5 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Wave quantization,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f69929ea-c697-4a01-94f8-042787a02f7c · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Nvidia fastertransformer: Transformer related optimization, including bert, gpt,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f84d62a-3cc6-4af2-a63b-164c31ab2b39 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ctranslate2,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f060c808-1a7c-4988-9be3-720d8f49e024 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Judging llm-as-a-judge with mt-bench and chatbot arena,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbe957d3-64d0-4562-ac55-ab52abffefe9 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLaMA: Open and Efficient Foundation Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985f7dad-d686-4f00-91bd-36f94c835bb3 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Training Verifiers to Solve Math Word Problems
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32f7167-717d-43fe-93a9-d99c0b8e39ff · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Evaluating Large Language Models Trained on Code
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaae6fc4-7c10-4305-896f-fc8f7963e2d3 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Program Synthesis with Large Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0de30b68-523e-419a-a089-cada98f314b7 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for CPU-GPU integrated edge devices,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c1462e8-01ac-4aa8-b1ee-df2badc6a558 · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c58228b5-db65-4718-9996-92174d32b9aa · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication-efficient model parallelism for distributed in-situ trans- former inference,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3743d6-38b0-495b-a278-4c1245722ebb · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Petals: Collaborative Inference and Fine-tuning of Large Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d670f62-08ce-4a51-8d72-5d4dec06a26e · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d0fb433f-8dc3-4c79-b65e-8b1912fdd96f · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4012f9de-80bc-48e5-bbbf-2660dc0bc53f · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLMCad: Fast and Scalable On-device Large Language Model Inference
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84bb3486-6f11-49d1-900f-b3d522f463cb · outbound
Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.