Pith. sign in

Paper Citation Record · LEDGER

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism

As of 8 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2505.23219.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23219 v2

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:02:10.198594Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact1
  • verified fuzzy21
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24ae3333-f6f1-45e2-97c8-e76e0b08e213 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Efficient memory management for large language model serving with pagedattention,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.403875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.403875Z digest=sha256:f3e7dd08b986fa8f1720864c3b841b4fd8f00030981c4c9f61c43c9a33f9a4eb

Observation 91206771-a7bc-4a05-a2bb-0b33cdef305d · outbound

This paper cites Specinfer: Accelerating large language model serving with tree-based speculative inference and ver- ification,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Specinfer: Accelerating large language model serving with tree-based speculative inference and ver- ification,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.911548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:06.537216Z digest=sha256:cd4b723ff80c54270c23fb7c7dfa65dea0c9c53219fdc77411caf09b5abd4b52

Observation 9fb6f691-5aac-4615-bf7b-817a44725dce · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.695859Z digest=sha256:7cd67a5bedefd7ebab63b4a598fbcb078c9cffedc63de85bb51f1d85c1c23ced

Observation 94627f43-b782-4c43-bcd0-91244600d07d · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Break the sequential dependency of LLM inference using lookahead decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.657708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:06.850459Z digest=sha256:c394c883026c5e443dfe181373c6cdd31eef637c29a30cc1d30283a464844d95

Observation a371d8a0-6efc-4a5c-9ca1-5897462139c6 · outbound

This paper cites Codl: efficient CPU-GPU co-execution for deep learning inference on mobile devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient CPU-GPU co-execution for deep learning inference on mobile devices,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:06.977455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:06.977455Z digest=sha256:0567cd51b7f82e22249ac2bb9096727924b69b7b5457ddf4342275f83867daf4

Observation 4e581d66-b47b-4ca8-b845-521ed67b90f7 · outbound

This paper cites Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for cpu-gpu integrated edge devices,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:16.396266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.117354Z digest=sha256:5e4f0a7f7950de39880d219bed7820653fb17825d09ac87633182cff796baec7

Observation 0c9e5fbb-7dfa-4666-9e70-05434b1d00ee · outbound

This paper cites High-throughput cnn inference on embedded arm big. little multicore processors,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism High-throughput cnn inference on embedded arm big. little multicore processors,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.960000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.252057Z digest=sha256:3ffbea15ff7e8b8e6e8a802f5f2b4b657db6b431ec0f10400458755fed0edad8

Observation 32fc7452-b213-4759-93c1-5844b13f9b00 · outbound

This paper cites Dopia: online parallelism management for integrated cpu/gpu archi- tectures,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Dopia: online parallelism management for integrated cpu/gpu archi- tectures,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.597486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.376808Z digest=sha256:1ee081f557f1d53b71d06207b76ccbc49a4710d0d0440e506cd8af23d8c29ff1

Observation c2baf9dd-8904-47ae-9ed1-bba66802decb · outbound

This paper cites Apple m4,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple m4,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:15.312135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.475644Z digest=sha256:d7299efec6445f437446ca59d0016bb0ba9259e1e4c9ff2cb240adb650f862d5

Observation a9dad62b-73f1-495f-8299-3ae6003b911f · outbound

This paper cites vllm github,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism vllm github,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.987334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.573662Z digest=sha256:f7b09d6a539a2ddb57edb7b108b09164294f958c4cbd9372c0762b4dd3b6e3f1

Observation 58b9bea6-9407-493a-97a4-732d593c3e7f · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:07.687263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:07.687263Z digest=sha256:07606bcb7bc3af39e46fe96e60de398f5420d63cb1eb19e360181eb8f7ceb705

Observation cefeb681-f9e5-4b0f-bc37-fde6ef58283c · outbound

This paper cites Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Galaxy: A Resource-Efficient Collaborative Edge AI System for In-situ Transformer Inference

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:07.770435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:07.770435Z digest=sha256:07b4d3321fcb299a9e51d1d85e6655bcf90ee2d3123a2f4586fc884cb2b746bb

Observation 7432d218-f45f-4fdd-8522-98a42ebbed96 · outbound

This paper cites Intel core ultra processor family,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Intel core ultra processor family,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.692481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.867375Z digest=sha256:138d080429a47d06507f90c90f47dc02115f46b9a16784e2e482de8f22dbd18d

Observation f3fc7d8d-41d0-46f8-9d9f-ff0a5f54c87e · outbound

This paper cites Meet jetson, the platform for ai at the edge,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Meet jetson, the platform for ai at the edge,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.344380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:07.993100Z digest=sha256:3e396b71570f86f83706fab9132b52c93f94448967636b4b1a92316d469e8c17

Observation f5589f97-86ef-487d-8f49-38482ae3dc32 · outbound

This paper cites Communication effi- cient distributed machine learning with the parameter server,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication effi- cient distributed machine learning with the parameter server,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:14.024404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.088083Z digest=sha256:2dc32669f69d2e4fa40e67f28f8503d842a466141e2b089cef5ffc374ce7ae49

Observation f016866a-0884-484b-9494-5f9d5be016f1 · outbound

This paper cites Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Deepspeed ulysses: System optimizations for enabling training of extreme long sequence transformer models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.716687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.196608Z digest=sha256:cba007c5d3fa3f0ea1c92ff2c365c1da7e39d6067c90ffc8e771cacb33e3dec8

Observation 5a4f5d34-2367-49d8-981a-1361d9009a5b · outbound

This paper cites Codl: efficient cpu-gpu co-execution for deep learning inference on mobile devices.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Codl: efficient cpu-gpu co-execution for deep learning inference on mobile devices

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.472004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.363551Z digest=sha256:f0c572d755eaadb4c661c596e7485ae67d06918ec92653e7ad4cc5555aa4f5d7

Observation 9c66d894-9c9a-4eb7-98b3-69b0b4c28a22 · outbound

This paper cites Attention is all you need,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Attention is all you need,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.445461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.445461Z digest=sha256:0d5016813af96fc666c9b44364846031884809ea2c2f5d3a98e782591054fc02

Observation 5bb6186d-b056-45b1-953a-ddd20e647f17 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.525545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.525545Z digest=sha256:f2445969a212a2dcdc2aeb4fdbe5f5d265536d57f6f559a9cc06b01424ff1b00

Observation f6096911-0308-49be-a5c8-4507020568d7 · outbound

This paper cites EAGLE: speculative sampling requires rethinking feature uncertainty,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism EAGLE: speculative sampling requires rethinking feature uncertainty,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:13.169659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.621537Z digest=sha256:86e08f4b0ffb08eb2914b1132c19d538a5f70d6c4efd4459ccbb937691a479cb

Observation 9f73d127-3366-4a33-8495-f3a5a873b60d · outbound

This paper cites Apple a17,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Apple a17,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.863170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.693852Z digest=sha256:8c15d92db027abb54a294185c72cd01649a87d42819743da54957ae7b9c8188f

Observation d9580a79-5031-479f-be40-09730113aa80 · outbound

This paper cites Amd reveals next-gen desktop processors for extreme pc gam- ing and creator performance,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Amd reveals next-gen desktop processors for extreme pc gam- ing and creator performance,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.552778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.789304Z digest=sha256:0935db0ce70996ee42638878b98d669f098d8946f41fa32fdc941a4a59d828cf

Observation 0a0d0a69-c945-4d6d-b371-6b32b9cb5ad6 · outbound

This paper cites Qualcomm snapdragon,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Qualcomm snapdragon,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:12.235562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:08.899326Z digest=sha256:0286bd4fb69b70c678243c2c5f3284f93b10f3e0574e511abc5ddc86f09d7ada

Observation 0ee5fff8-57fe-4f4d-8b57-00257fc7c042 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.016191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.016191Z digest=sha256:98b09dff7a59be627401de4c32139117d641939cc6a8c5b3eb94b10ee2761e33

Observation 29985fa0-3612-487d-8f4e-3f1164b86d7e · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.920095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.116736Z digest=sha256:6b151786a7b48a60a8fa73e73cf652eff6b7deeb005fe08eafa5e1ed72bd3930

Observation 9b6bcb50-3f38-4bdc-9901-6ba9cc0e51d5 · outbound

This paper cites Wave quantization,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Wave quantization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.668637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.228587Z digest=sha256:6bfb3a9c3770ceb680709456489038ff050dbc31eb97301c393bc42a0fc1b62c

Observation f69929ea-c697-4a01-94f8-042787a02f7c · outbound

This paper cites Nvidia fastertransformer: Transformer related optimization, including bert, gpt,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Nvidia fastertransformer: Transformer related optimization, including bert, gpt,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.430108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.390790Z digest=sha256:ee7652658612e72f0769b0cfbef0a97072b1b9ca7a58d650cbb9083149f17b61

Observation 9f84d62a-3cc6-4af2-a63b-164c31ab2b39 · outbound

This paper cites Ctranslate2,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Ctranslate2,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.228226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.458408Z digest=sha256:0202e6f073c3667c919fc2a0523c98b44371276d9c8ab89488975a4b39f27aa2

Observation f060c808-1a7c-4988-9be3-720d8f49e024 · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Judging llm-as-a-judge with mt-bench and chatbot arena,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.524187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.524187Z digest=sha256:8deb80d37087e7cb224f879a21aa8e20927732bf52f6d435ed3e7102c576963a

Observation cbe957d3-64d0-4562-ac55-ab52abffefe9 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLaMA: Open and Efficient Foundation Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.585286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.585286Z digest=sha256:b849b0688795564e16a88b31a4716d3f3d2f3dc53db857cbbb39427ca1d3669d

Observation 985f7dad-d686-4f00-91bd-36f94c835bb3 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Training Verifiers to Solve Math Word Problems

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.637595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.637595Z digest=sha256:efcfd582d1f865e827d18f8a0a5a552e6d14436123d7f9e0dbaee8baa079a3c3

Observation b32f7167-717d-43fe-93a9-d99c0b8e39ff · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Evaluating Large Language Models Trained on Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.703036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.703036Z digest=sha256:d7cbc4ae49890b80d604e3759ae8ae72c921f4b7f7a28a8d9c3d8837f6b06d68

Observation eaae6fc4-7c10-4305-896f-fc8f7963e2d3 · outbound

This paper cites Program Synthesis with Large Language Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Program Synthesis with Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.762157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.762157Z digest=sha256:c086ca7908b047e82149d4c40b35fd007a445448c8a79c6c848866bfd90fc926

Observation 0de30b68-523e-419a-a089-cada98f314b7 · outbound

This paper cites Edgenn: Efficient neural network inference for CPU-GPU integrated edge devices,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Edgenn: Efficient neural network inference for CPU-GPU integrated edge devices,

Reference 34

Resolution
verified exact
raw_fallback, observed 2026-08-07T13:02:10.643834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.831626Z digest=sha256:dba553536ecc6ab0472a702d0ea61094124260829fcdde4db7645158f910d158

Observation 5c1462e8-01ac-4aa8-b1ee-df2badc6a558 · outbound

This paper cites Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Taming {Throughput-Latency} tradeoff in {LLM} inference with {Sarathi-Serve},

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:02:11.067511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:09.895407Z digest=sha256:e53d77563cef0be7de9a3be3c310062398212ba5d5958c788abd3b9bff1df46d

Observation c58228b5-db65-4718-9996-92174d32b9aa · outbound

This paper cites Communication-efficient model parallelism for distributed in-situ trans- former inference,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Communication-efficient model parallelism for distributed in-situ trans- former inference,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:09.955947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:09.955947Z digest=sha256:be30f38fea8a1e2aaa4c5a330aaca014c825d25c5ad4ca817f1cad9fe28a87cf

Observation 3d3743d6-38b0-495b-a278-4c1245722ebb · outbound

This paper cites Petals: Collaborative Inference and Fine-tuning of Large Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Petals: Collaborative Inference and Fine-tuning of Large Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.010089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.010089Z digest=sha256:3eae57993e8ce2a0945563992528cbb621f8cec698907acb210baf2aa15f9517

Observation 0d670f62-08ce-4a51-8d72-5d4dec06a26e · outbound

This paper cites Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus,.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Asymo: scalable and efficient deep-learning inference on asymmetric mobile cpus,

Reference 38

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T13:02:10.429864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T13:02:10.076767Z digest=sha256:8292e90285e70a98f28ba7fcdfd0570251868e4e425392d33747c405ebcdf5c8

Observation d0fb433f-8dc3-4c79-b65e-8b1912fdd96f · outbound

This paper cites Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism Sequoia: Scalable, Robust, and Hardware-aware Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.146184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.146184Z digest=sha256:717df7fb51d53ef2f482fededc1232d4e5d4395a25d5a5aa6ee989f6b04196dc

Observation 4012f9de-80bc-48e5-bbbf-2660dc0bc53f · outbound

This paper cites LLMCad: Fast and Scalable On-device Large Language Model Inference.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism LLMCad: Fast and Scalable On-device Large Language Model Inference

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:10.198594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:10.198594Z digest=sha256:494415a44b593e6eedc3bb94f6094b5cfbf9f524dfca80931f9b95a7bc78892d

Observation 84bb3486-6f11-49d1-900f-b3d522f463cb · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Ghidorah: Fast LLM Inference on Edge with Speculative Decoding and Hetero-Core Parallelism DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T13:02:08.290303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:02:08.290303Z digest=sha256:222cd8412d7ed6f6af3f60cb579e2e010da7f01ab72cd17d77c57072a1951972

Pith citing papers

No inbound Pith citation observations are available.