Pith. sign in

Paper Citation Record · LEDGER

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

As of 21 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2506.09554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09554 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:49:05.158865Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:58.959186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.580222Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8af69b5d-aa14-4aa8-a383-048688a79ac6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.534373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.534373Z digest=sha256:ddf86214acd02f659032e6b2c31be16cc332fec6625a72a15490a6e1bbfe08a9

Observation 9f5b5f4c-0136-46ec-8aab-02981a2b5328 · outbound

This paper cites Plug in the safety chip: Enforcing constraints for llm-driven robot agents,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Plug in the safety chip: Enforcing constraints for llm-driven robot agents,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.996503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:03.625036Z digest=sha256:1edee8db13612e3b982ae5df0dd387b904d67b1bb592e653531b6a7fabb9d5dc

Observation 620bec5f-18e1-4327-a0c7-fcc3e9d3281c · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A Survey on Efficient Inference for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.789514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.789514Z digest=sha256:21ac06a0ed284e16f8837dfc8000692bee2e5bb8161fdd11c4ae037b27fca5c3

Observation 89ecc77c-05b5-476d-8f20-896aca0967b4 · outbound

This paper cites LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.916941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.916941Z digest=sha256:34ec3243d4e8ae3ea85e8290ad71f51e9d56437f80e6055ac0d3a59ffa25c741

Observation b6633e4d-7b70-499f-83b9-4d4bc9029cc4 · outbound

This paper cites Characterizing the performance of accelerated jetson edge devices for training dnns,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Characterizing the performance of accelerated jetson edge devices for training dnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.859489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.079899Z digest=sha256:2db5564279060433cc155c84c09a7723d8657f8eeb3e20d1f33211ed83c70cce

Observation 550cf011-0485-43ba-8ed4-d1302fef8095 · outbound

This paper cites Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:49:05.357455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.191036Z digest=sha256:f6b8e3a743599fb37616b151ac7dfba883ee2156dbffc01bbb69f4c4d857460d

Observation 5952f333-940e-4bb0-8077-efa06d8ea389 · outbound

This paper cites A preliminary performance analysis of llm inference on edge accelerators,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A preliminary performance analysis of llm inference on edge accelerators,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.675928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.296823Z digest=sha256:89c87d35b8b47c612d367b37c2b2948aecb738cfe23895a21694864d87ee4276

Observation 90bfe4cf-22eb-4e00-989f-2a49948496f8 · outbound

This paper cites Wikitext-2,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Wikitext-2,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.518047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.415311Z digest=sha256:936408bd9c3c2fd8210c4322206f9c1d6242f3ce73422e8ca9be1ade265e776f

Observation 32235912-1aff-4d2f-b2b7-0818a9672b08 · outbound

This paper cites Longbench,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Longbench,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.334059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.556604Z digest=sha256:d8a2af7f7434fc08a0691ee4f34b6e58e28d74092bbd76306758ffdb99f6a128

Observation 37ed37ed-a5e6-4ed3-a6ea-ef8ce45a6286 · outbound

This paper cites Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.160849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.676505Z digest=sha256:60ca9596f063126e22b2b02846dc2c017a4230beb89f77ea5e757670303577d8

Observation 8315a6d7-71c0-4c55-ab46-ddabf1f332ec · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Splitwise: Efficient generative LLM inference using phase splitting

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:49:04.761986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:04.761986Z digest=sha256:ffba8825a6ed8af4dc4986f4afe85c61fe81e3096426ef761bf265a716fc7f02

Observation 0f1a1851-5edd-4ea6-b0f6-8c6664812145 · outbound

This paper cites When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%).

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%)

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:49:06.022150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.852213Z digest=sha256:23c228935203a421345068187ebc73ddfd1a070ba50b2eee1b3c0bca78b8a687

Observation cc80d7e8-585a-4c02-b903-569a8cb2d6b3 · outbound

This paper cites Against INT4, INT8 consistently yields over 27% power savings.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Against INT4, INT8 consistently yields over 27% power savings

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.886948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:04.987410Z digest=sha256:6b6bd9ef5abc55c3e93cd9e0ed0dbbf200093d420c39ff140829d8860a29048a

Observation cb242ee6-3e20-4f56-bead-44e9b01bc9a0 · outbound

This paper cites •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.737086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:05.068187Z digest=sha256:ed2444368c0187ff6fb0db028b64a26c38b61b98d9ef3f8429e0c2cc10f25373

Observation dfb5deec-c2a4-4ee6-b37a-dce61653e0fb · outbound

This paper cites an unresolved cited work.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:05.557627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T04:49:05.158865Z digest=sha256:6d4576c17daea7844a47baee6fa74dab0265f877fd7bdbf6e276e588066e4b5c

Pith citing papers

Observation 09e339dd-26dd-4613-8d3b-302929859ad7 · inbound

On the Sustainability of AI Inferences in the Edge cites this paper.

On the Sustainability of AI Inferences in the Edge Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:58.959186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:58.959186Z digest=sha256:331f4ee7b9860802a3567b08f8cbae5804e7c45aefb85dccd983853dc32073b0

Observation 056f7926-2373-437c-93e6-ce0d485aa266 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.582044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:d01dc816339beec7ba879ecf7f5c950cc039ec0e4946d8672f946cfb3e53abbd

Observation bfb30276-44c1-4ed4-866f-de8350740453 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:1f4d7e3812243a33860c4ed8fe8eee891e883d8fd98d6f5571098e245d05a69b