Pith. sign in

Paper Citation Record · LEDGER

Large Language Model Partitioning for Low-Latency Inference at the Edge

As of 17 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2505.02533.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02533 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:54:46.812096Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-08T09:45:57.201837Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T20:16:10.526836Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a272bea-1e7c-4b8e-bc43-2cb9a607dc3e · outbound

This paper cites EdgeShard: Efficient LLM Inference via Collaborative Edge Computing.

Large Language Model Partitioning for Low-Latency Inference at the Edge EdgeShard: Efficient LLM Inference via Collaborative Edge Computing

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:46.730815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:46.730815Z digest=sha256:32138dfad82d53447614f15f5006bc16a82b182e5062d6835f8008279802890f

Observation cf602b55-b084-4bb5-8deb-4cd5c7136544 · outbound

This paper cites SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization.

Large Language Model Partitioning for Low-Latency Inference at the Edge SplitLLM: Collaborative Inference of LLMs for Model Placement and Throughput Optimization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:46.736406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:46.736406Z digest=sha256:8956c0c4552888fe120976b85d89df6900ddec69ba1fd2c31c7139cc0405e17f

Observation 0a82c498-8154-4fba-85de-ccc1754349e3 · outbound

This paper cites Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.114469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.741716Z digest=sha256:609092862a27126a1516a017a46532473b84e3984fe59aae2fd4fe68a61b1978

Observation ac62089e-ffc0-4aed-a02c-a83caae85127 · outbound

This paper cites Language mod- els are few-shot learners,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Language mod- els are few-shot learners,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.102495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.746247Z digest=sha256:91760491df8a18e032131a3ecc4072293606d242f537bc2b305434ee6fb49108

Observation e0edf7a8-d28a-45ea-93bb-41b95800e1f1 · outbound

This paper cites 6g technology overview,.

Large Language Model Partitioning for Low-Latency Inference at the Edge 6g technology overview,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.088538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.752903Z digest=sha256:9e78808d98796c86b7aaaa27d5fbd0e8991d7b7757e3d4b448348c5a9b76a83a

Observation 5256ce74-4766-410f-b1c0-838730c2dac3 · outbound

This paper cites Dnn partitioning and inference task offloading in 6g resource-constrained networks,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Dnn partitioning and inference task offloading in 6g resource-constrained networks,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.076062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.756877Z digest=sha256:6a9ae44abd0f1f5025de6e25dfd396a4f52f27da1f5d6e143d7ad64cb110e3ac

Observation f2541d85-f014-479b-a75f-e832e283a9e8 · outbound

This paper cites Splitplace: Ai augmented splitting and placement of large-scale neural networks in mobile edge environments,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Splitplace: Ai augmented splitting and placement of large-scale neural networks in mobile edge environments,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.063083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.760997Z digest=sha256:e4179fb3058895962bb75f105611833ad6a0b8560339a1d1b54d7ded97926b30

Observation 769c876a-f3ff-4c81-92c2-2d9ea2300583 · outbound

This paper cites Joint optimization of the partition and scheduling of dnn tasks in computing and network convergence,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Joint optimization of the partition and scheduling of dnn tasks in computing and network convergence,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.049461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.765023Z digest=sha256:9019cb465f6585f1b0bbbaa9ea98df1e98dd1ee235b35544450cc03f44dfb590

Observation 3f934288-8d4d-40db-a7c1-b8df6506c1fe · outbound

This paper cites Joint optimization of dnn partition and continuous task scheduling for digital twin-aided mec network with deep reinforcement learning,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Joint optimization of dnn partition and continuous task scheduling for digital twin-aided mec network with deep reinforcement learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.035460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.768989Z digest=sha256:613174d3157627a9aead83e47d15af1392ff9289da1e51ecd1f365f2defc45ea

Observation 64a78c78-37e5-43fc-9dde-006eb77ebf0c · outbound

This paper cites Huang, Y.

Large Language Model Partitioning for Low-Latency Inference at the Edge Huang, Y

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:46.772434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:46.772434Z digest=sha256:f3496a37de8eb22e2e3460cd0138af9d0475c9dd77e81899a285452d2601d8bf

Observation 3391209d-cded-4304-821b-27f1cd320ebd · outbound

This paper cites Pipedream: generalized pipeline parallelism for dnn training,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Pipedream: generalized pipeline parallelism for dnn training,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:47.011289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.776387Z digest=sha256:23d1df4fddc859a7efa6e41b19a24e6a974f164162cd0246f848c95fab707455

Observation c0667170-2416-4641-b0ab-8d515958be94 · outbound

This paper cites Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Autopipe: A fast pipeline parallelism approach with balanced partitioning and micro- batch slicing,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:46.977365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.780076Z digest=sha256:9aa6f561bc6557891a804bfa63ef5721b9d74f427efe6f727aad0fc729146e20

Observation d2f0ab22-8510-4a26-92a6-ec7f821d95c9 · outbound

This paper cites Memory-efficient pipeline-parallel dnn training,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Memory-efficient pipeline-parallel dnn training,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:46.946864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.784317Z digest=sha256:245055da05f107fee4519a5f5b207a80cf35e72b3042d4171113b3a595064e78

Observation 1d614240-fd22-4cae-a086-59489a1db3b5 · outbound

This paper cites 3d parallelism for transformers via integer programming,.

Large Language Model Partitioning for Low-Latency Inference at the Edge 3d parallelism for transformers via integer programming,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:46.916654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.788881Z digest=sha256:9881fbd03a0551374a8c25b3c98fd609d45c228dea4f444823a3563d5dc29dcc

Observation 457e8af3-24fa-402b-b136-fa2a9b3e526d · outbound

This paper cites Distributed Transformer Inference Simulator.

Large Language Model Partitioning for Low-Latency Inference at the Edge Distributed Transformer Inference Simulator

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:46.897800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.800786Z digest=sha256:8b58f25c98fbeec424b4a0b6272d2acb23e87071ae27325de9f52c20efce67d8

Observation f32dd22c-79ea-43f2-8aaf-30b4c766e4cc · outbound

This paper cites Google cluster-usage traces: format+ schema,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Google cluster-usage traces: format+ schema,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:54:46.883972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T00:54:46.808123Z digest=sha256:798ead2578978878365df383b6e574c06cdbbe42eaf7dc17527ea492a553e1f3

Observation 4c9939d6-f98b-4cfa-831d-019d7571081a · outbound

This paper cites Attention is all you need,.

Large Language Model Partitioning for Low-Latency Inference at the Edge Attention is all you need,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:54:46.812096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:54:46.812096Z digest=sha256:db7b1f70f091103ddc9aa8b235dd596c33e5885709cce52b9847c1ade4a3389e

Pith citing papers

Observation 16d738a7-7593-408d-b3bb-463e39467121 · inbound

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities cites this paper.

Network Edge Inference for Large Language Models: Principles, Techniques, and Opportunities Large Language Model Partitioning for Low-Latency Inference at the Edge

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:10.529259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T09:45:57.201837Z digest=sha256:3d5afe26e27a741bb4c046748367a878cdfb6fca7f94a9a537b7887885c86c49