Pith. sign in

Paper Citation Record · LEDGER

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding

As of 7 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 0 inbound Pith citation observations for arXiv:2508.12590.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.12590 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:31:39.052107Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13ffb03c-64ba-4c7f-8177-5843e13104ef · outbound

This paper cites Harnessing the power of llms in practice: A survey on chatgpt and beyond,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Harnessing the power of llms in practice: A survey on chatgpt and beyond,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.461160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.002774Z digest=sha256:daec3a113a78a17aacf5af9dd7bc2e80b932652e785fc5e66466bd0b6f1e678d

Observation e68eeda1-f9d8-4850-a7e5-c2f12fb57fdd · outbound

This paper cites Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Pushing Large Language Models to the 6G Edge: Vision, Challenges, and Opportunities

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.005967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.005967Z digest=sha256:95d14ffc9f9051201de2e9989c3e45d21bcc10e4faadb8e954a1c696834e8a6c

Observation 90689d9d-85e8-4f72-83f0-02512697cc30 · outbound

This paper cites Hybrid slm and llm for edge-cloud collaborative inference,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Hybrid slm and llm for edge-cloud collaborative inference,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.451366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.009391Z digest=sha256:3d944f75def3286c2d5fdc90b46c2635db51e2d569fa4bec60c0701c49bb1a31

Observation fbdc6fd9-d6e9-4907-ab36-08da220991de · outbound

This paper cites Fast inference from trans- formers via speculative decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Fast inference from trans- formers via speculative decoding,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.441705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.012982Z digest=sha256:75dfa3b1a4c7252f9ac10a235c54ced6d843e90224829826d3ed3cbf669fcdbf

Observation 81bbfd8e-ab94-4f7e-8c4a-830767f993fe · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.016413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.016413Z digest=sha256:e5be733b058efa9b618c4c062d795fee3ef18965904a766c19a7776bdc05277a

Observation 4d0d7295-cda8-4254-a264-dfb25f134a61 · outbound

This paper cites Uncertainty-aware hybrid inference with on-device small and remote large language models,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware hybrid inference with on-device small and remote large language models,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.431631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.019737Z digest=sha256:ea16214401e939781620ddbfc3085abf371a60d5b62e77cb1c5c5dc850153b6c

Observation dbfd4561-579e-4f76-b464-79dd46aef187 · outbound

This paper cites Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Uncertainty-aware opportunistic hybrid language model in wireless robotic systems,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.411228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.022856Z digest=sha256:5395eedc2efc50730ab5e562c8d843040d9318b8d2929e723e70e57d71d98c11

Observation f0b85060-eb4e-4bc8-9d85-93284ee55516 · outbound

This paper cites Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Seeing far and clearly: Mitigating halluci- nations in mllms with attention causal decoding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.397924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.025859Z digest=sha256:d71e84be6cef6ba0081c88382489b09b10e2edfdec1252e0fc2298741e22c05d

Observation 3c9b5960-3ac4-42fd-9706-84f200110ca4 · outbound

This paper cites Understanding the metropolis-hastings algorithm,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Understanding the metropolis-hastings algorithm,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.378874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.028602Z digest=sha256:839ec98bd7e8f97aa6dfc6574ce470fd0f610038b8b64db8cc5881f6fd5894fe

Observation 130c4ab6-96f5-4c1c-a11f-e393be730b73 · outbound

This paper cites Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Attention Score is not All You Need for Token Importance Indicator in KV Cache Reduction: Value Also Matters

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.031313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.031313Z digest=sha256:c1cde36e8cd3e66715be90f765e73ee74e13bd423281006c7c77ba6f46981144

Observation 9b3dc1b0-98b5-458e-a22b-c7ebd21d54ec · outbound

This paper cites Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Latte: Low-precision approximate attention with head-wise trainable threshold for efficient transformer,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.341197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.034625Z digest=sha256:61a770c316502a863d17efe4bc0fbb3433d0a7ae2a339291e307c8ba6c83a5e8

Observation afd689fe-6809-4571-98fd-d42752e95ac7 · outbound

This paper cites Energon: Toward efficient acceleration of transformers using dynamic sparse attention,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Energon: Toward efficient acceleration of transformers using dynamic sparse attention,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.259478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.037367Z digest=sha256:d1759efad42c33317b91514e0e7914b909799aba092f6d29f59108b813c7c9e4

Observation fae26586-119d-4841-ba5f-6bd02c280899 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding TinyLlama: An Open-Source Small Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.039834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.039834Z digest=sha256:5a911a19755e8ffa1dd11ab2b390b45b84b2e7dc5d5a725b9b186f651d88604b

Observation 079bf107-d6d2-426b-9f3b-5d8852fbc3a6 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.043735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.043735Z digest=sha256:db63fe11d9b587e101b133943b76d22109cd7823519e27eb2fe8043131605d0b

Observation 966da510-63e7-4b40-a262-12ba3c6448d5 · outbound

This paper cites Stanford alpaca: An instruction-following llama model,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Stanford alpaca: An instruction-following llama model,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.046483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.046483Z digest=sha256:8c17bd32571b8d24b9064983c71c9a166f417b4fe09ab9f8cb8a7ad7c091e65b

Observation ff3167e8-2ec7-4f99-bb73-feafd452aaf7 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding BERTScore: Evaluating Text Generation with BERT

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:31:39.049209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:31:39.049209Z digest=sha256:1f1f56b92ad7ceee9c2924b8e5c8d92a4c5a29bf43bbd70b159cbf8d7c7e912d

Observation 36761e8f-6eb2-4193-8387-555625992248 · outbound

This paper cites Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,.

Energy-Efficient Wireless LLM Inference via Uncertainty and Importance-Aware Speculative Decoding Joint computation and communication cooperation for energy-efficient mobile edge comput- ing,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:31:39.162658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T19:31:39.052107Z digest=sha256:581773f09cbae9d482543ebd8ea6743074296cd1687a38c4dc14d4c788203272

Pith citing papers

No inbound Pith citation observations are available.