Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:04:37.517635Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2512.21835.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T14:04:37.517635Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 347964fc-61d7-4a83-a328-98418ecb40bf · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices DeepSeek-V3 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f067646-16b2-4be3-8163-2e9422b6aab5 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Improving language understanding by generative pre-training,
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07d77fe5-ce41-4b45-a2df-b5a7371ec808 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sasha: creative goal-oriented reasoning in smart homes with large language models,
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec56baf-10b9-4b35-8efb-e63026edc1ae · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Gemini Robotics: Bringing AI into the Physical World
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1158706-503f-43b3-a1c1-f6546f429006 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Privacy Inference Attacks and Defenses in Cloud-based Deep Neural Network: A Survey
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09d28748-6dee-4d5b-8e0f-faa145f3da81 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Privacy-preserving federated learning for transportation mode prediction based on personal mobility data,
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9663b91e-ecee-461f-ace8-78b01e93401a · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Re- search on medical data storage and sharing model based on blockchain,
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5b9380-7a76-4f85-9f18-62f21d860edf · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices {ServerlessLLM}:{Low-Latency}serverless inference for large language models,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a95d7e2-e5e6-4848-b422-fa5bd7b96aa3 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices The llama 3 herd of models,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 08d95cc8-6c3c-4188-a13b-991ba6520659 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Nvidia jetson xavier,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bce7fa31-24ac-4b45-979c-dddb123b5dbb · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Kvquant: Towards 10 million context length llm inference with kv cache quantization,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c40f42f-b3c2-42b6-ac17-82841a3ccebe · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices SpQR: A Sparse-Quantized Representation for Near-Lossless LLM Weight Compression
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 65590f5b-6841-4d13-a596-b45f86c62864 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Qlora: Efficient finetuning of quantized llms,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1222fb62-48b8-4dcb-9636-2ab067d195f6 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 107578cc-9dea-4a79-8380-29dc03b4f203 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f627b4a-b326-4e4e-bfb1-8f4a4da11880 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices MiniLLM: On-Policy Distillation of Large Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aad829c-bdec-4f0c-bc3e-20f7b6c5e14e · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Less is more: Task-aware layer-wise distillation for language model compression,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe458615-5a4c-4155-8bd1-677c24c9fcaf · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Distilling step-by-step! outperforming larger language models with less training data and smaller model sizes,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e30c8108-2404-47a1-849f-5cd8af5bafd9 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Resource-aware federated self-supervised learn- ing with global class representations,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6868cb3d-a0ed-43ad-8b16-8f1ed472c0f7 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sparsegpt: Massive language models can be accurately pruned in one-shot,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 426fa491-f22c-41a2-a4c0-9414821500d4 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Unity is power: Semi-asynchronous collaborative training of large-scale models with structured pruning in resource-limited clients,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6bee5a92-c87e-4a35-89dd-8fa390fecfbf · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices SliceGPT: Compress Large Language Models by Deleting Rows and Columns
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6634eded-36bc-4e47-b5fd-8f02d0ac0ebc · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Llm-pruner: On the structural pruning of large language models,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 276674d8-e370-4885-8b8b-a7961b363594 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Generalization-aware distributed minimax optimization for large-scale models on resource-limited devices,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5284734-b987-4907-9914-2d0fc84c0a64 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Theoretical convergence guaranteed resource-adaptive federated learning with mixed heterogeneity,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9348d74-8ef6-4472-b5c2-f4051b35c602 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Credit risk analysis using machine and deep learning models,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eab24cbf-13ad-42b4-b6dc-79d22740a56a · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Medical image analysis using deep learning algorithms,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28cfe570-720f-4106-830b-2767489d3dd0 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0336e8dd-c341-4c79-9508-165368199780 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Edgeshard: Efficient llm inference via collaborative edge computing,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77669c8c-c35c-431b-8035-9531a73be352 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Galaxy: A resource-efficient collaborative edge ai system for in-situ transformer inference,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28889b35-fdbd-408b-b30a-7d1d1259c442 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Tpi-llm: Serving 70b-scale llms efficiently on low-resource mobile devices,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f461e3b-507e-4f4b-83d4-94c9134b0a53 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Band: coordinated multi-dnn inference on heterogeneous mobile processors,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8fb5fd0-5b98-40a6-935e-8707d523343b · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Coedge: Cooperative dnn inference with adaptive workload partitioning over heterogeneous edge devices,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9d5f377-fc33-4055-b2df-052b3ab73c85 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Model parallelism optimization for distributed inference via decoupled cnn structure,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25b0704f-d791-4f6b-b628-100d4179953b · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices When the edge meets transformers: Distributed inference with transformer models,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416b80cd-ab38-49ad-ae6c-0d35398c08e0 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Jupiter: Fast and resource-efficient collaborative inference of generative llms on edge devices,
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 076bfe50-5ba5-4d3a-b77c-fb6b595a5533 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Sequence parallelism: Long sequence training from system perspective,
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 748c3806-7c81-4d1d-b196-73247132d685 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b052bd9a-b3da-46cd-9cb8-bac4877d8eae · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Pipedream: Generalized pipeline parallelism for dnn training,
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 399f08c9-c621-4d54-9d73-0928a71ac8d6 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Gpipe: Efficient training of giant neu- ral networks using pipeline parallelism,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3df890f-a88d-4903-a1f1-4aef2fa3fd01 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ba45872c-6f51-40ae-bda4-b3e3d980f3ab · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Flexgen: High-throughput generative inference of large language models with a single gpu,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 838c0edd-4585-4dcc-bfac-e0a520bf3eb4 · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Linux advanced routing & traffic control howto,
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a6f689e-e879-40b9-a078-8d8160822cbf · outbound
Collaborative Lossless LLM Inference Serving with Offloading-based Pipeline Parallelism on Edge Devices Qwen3 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.