Pith. sign in

Paper Citation Record · LEDGER

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 3 inbound Pith citation observations for arXiv:2506.09554.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09554 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:49:05.158865Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:58.959186Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T07:59:39.580222Z

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy8
  • unresolved4
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8af69b5d-aa14-4aa8-a383-048688a79ac6 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.534373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.534373Z digest=sha256:ef8386bbbceda18edd865193ff81c5b535a7de9cb2880c0556f24ac9bc1e3b9e

Observation 9f5b5f4c-0136-46ec-8aab-02981a2b5328 · outbound

This paper cites Plug in the safety chip: Enforcing constraints for llm-driven robot agents,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Plug in the safety chip: Enforcing constraints for llm-driven robot agents,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.996503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:03.625036Z digest=sha256:4ba53e27d669cc64323557650d28a0d755c6bb4f38a32b7047fa949db7ae7a41

Observation 620bec5f-18e1-4327-a0c7-fcc3e9d3281c · outbound

This paper cites A Survey on Efficient Inference for Large Language Models.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A Survey on Efficient Inference for Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.789514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.789514Z digest=sha256:7acc3977a26e1a5ccd3a87f24c4d4b236bcc7eebc34674cbd26092f720f5bc05

Observation 89ecc77c-05b5-476d-8f20-896aca0967b4 · outbound

This paper cites LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators LLM-Inference-Bench: Inference Benchmarking of Large Language Models on AI Accelerators

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:49:03.916941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:03.916941Z digest=sha256:0ecfba767b37f1f18482259e018ce9e7a5508ca2a5ab3ade8d7a515006b33697

Observation b6633e4d-7b70-499f-83b9-4d4bc9029cc4 · outbound

This paper cites Characterizing the performance of accelerated jetson edge devices for training dnns,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Characterizing the performance of accelerated jetson edge devices for training dnns,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.859489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.079899Z digest=sha256:28f2a45a5e34f44fb4280569d90d721ce5e45e70d5722ef945c2b1d26132dca8

Observation 550cf011-0485-43ba-8ed4-d1302fef8095 · outbound

This paper cites Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Large Language Models on Small Resource-Constrained Systems: Performance Characterization, Analysis and Trade-offs

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:49:05.357455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.191036Z digest=sha256:1d6825dc9689b3ddda29dee55085aa0698a3cc2caa905f0eb2908246071623ac

Observation 5952f333-940e-4bb0-8077-efa06d8ea389 · outbound

This paper cites A preliminary performance analysis of llm inference on edge accelerators,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators A preliminary performance analysis of llm inference on edge accelerators,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.675928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.296823Z digest=sha256:df75debdaf38b65bc7e9cbed3f5ba41c3037f133e797b6164d393dbb98065e07

Observation 90bfe4cf-22eb-4e00-989f-2a49948496f8 · outbound

This paper cites Wikitext-2,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Wikitext-2,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.518047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.415311Z digest=sha256:dd6870aa6ab26634f085e968f226d6f5585c3f9ad52f2a9f563a8e838959b735

Observation 32235912-1aff-4d2f-b2b7-0818a9672b08 · outbound

This paper cites Longbench,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Longbench,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.334059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.556604Z digest=sha256:49fc933dae26b379ccf4e91ebd1b1b6faf1304909b29f62e9d96e06641fa3aaf

Observation 37ed37ed-a5e6-4ed3-a6ea-ef8ce45a6286 · outbound

This paper cites Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Gpt3.int8(): 8-bit matrix multiplication for transformers at scale,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:06.160849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.676505Z digest=sha256:368084c4c4c4128d5697560679558fd3a2829c29101efb180d5809afbf339e8c

Observation 8315a6d7-71c0-4c55-ab46-ddabf1f332ec · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Splitwise: Efficient generative LLM inference using phase splitting

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-07T04:49:04.761986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:49:04.761986Z digest=sha256:73862d0fe25bdcc7b994c3894728b1520d71b997de03d816d77df97222bc265c

Observation 0f1a1851-5edd-4ea6-b0f6-8c6664812145 · outbound

This paper cites When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%).

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators When compared with INT4, INT8’s power savings range between 20% and 43% (median 32%)

Reference 12

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T04:49:06.022150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.852213Z digest=sha256:e7a02873afd6903e33e78f82dbbe67a795958a9cbdc4ebf65c720067861607d0

Observation cc80d7e8-585a-4c02-b903-569a8cb2d6b3 · outbound

This paper cites Against INT4, INT8 consistently yields over 27% power savings.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Against INT4, INT8 consistently yields over 27% power savings

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.886948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:04.987410Z digest=sha256:dfd813dde8c7124ca21756e2f830ec098b525fd18e64cc85f14503faea5bdaef

Observation cb242ee6-3e20-4f56-bead-44e9b01bc9a0 · outbound

This paper cites •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators •Energy Consumption: Similarly, INT8 achieves lower energy usage, with a median reduction of 24% compared to FP16 and 55% com- pared to INT4

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:49:05.737086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:05.068187Z digest=sha256:828947ad59d48827dc02d8649fd3e4a029cd48d3fc9f138a8882a3b59e7e03d3

Observation dfb5deec-c2a4-4ee6-b37a-dce61653e0fb · outbound

This paper cites an unresolved cited work.

Understanding the Performance and Power of LLM Inferencing on Edge Accelerators Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:49:05.557627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T04:49:05.158865Z digest=sha256:3c574e32ef696b9fcb88de8b5eb70a722e355e94d7d111cc6cec1ed3d4b20869

Pith citing papers

Observation 09e339dd-26dd-4613-8d3b-302929859ad7 · inbound

On the Sustainability of AI Inferences in the Edge cites this paper.

On the Sustainability of AI Inferences in the Edge Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:58.959186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:58.959186Z digest=sha256:bf36471c7856143bc202c081748819741e2bd35b26d3a4abfdd81bee8bd7c6ef

Observation 056f7926-2373-437c-93e6-ce0d485aa266 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T07:59:39.582044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-26T12:30:55.628115Z digest=sha256:a94507a5cb58d67b5e4895995942fcff71719db092cf636e4f4f0f3045408121

Observation bfb30276-44c1-4ed4-866f-de8350740453 · inbound

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study cites this paper.

Does Mixture-of-Experts Actually Help Inference on Consumer and Edge Hardware? An Empirical Study Understanding the Performance and Power of LLM Inferencing on Edge Accelerators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T13:05:17.273287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T13:05:17.273287Z digest=sha256:eff220e2a62e2a30334d33aadc424610f6d0ce4eb7737671eaf354a747368e45