Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:42:19.865695Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.13093.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:42:19.865695Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 704996cd-c2f2-4eff-93fe-991f9522c05a · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a244b28b-a60f-4a5a-8424-34b6eba53ff3 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbf50e4c-e461-40a9-8de9-8898dd6a5caa · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Flashattention-2: Faster attention with better parallelism and work partitioning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59aafbdb-ea91-4138-abbc-039fc7e78970 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ORCA: A distributed serving system for transformer-based generative models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ec4aa2a-05b0-4e4d-9fa1-fbca4f3b12e6 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a3237f2-631f-4397-abe2-2bd07717c460 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Neurosurgeon: Collaborative intelligence between the cloud and mobile edge
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c8ebfc-c67a-4f6c-8050-c791406b6084 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split computing and early exiting for deep learning applications: Survey and research challenges.ACM Computing Surveys, 2022
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be4ee5b6-3c18-4bc6-ace7-ce68744ba00a · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DNN surgery: Accelerating DNN inference on the edge through layer partitioning.IEEE Transactions on Parallel and Distributed Systems, 2023
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a7f8188-0c65-4f6f-be70-bde979f3536d · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3e7d1d7-cc21-4428-a652-4b0f664a237c · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hybrid SLM and LLM for edge-cloud collaborative inference
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33614f3b-2045-4f8a-b6b4-4898256a8cef · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Jupiter: Fast and resource-efficient collaborative inference of generative LLMs on edge devices
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50603c84-6301-4705-b54b-b447c97d11af · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbd50d06-12a4-4179-8465-1e9474670daa · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models CE-CoLLM: Efficient and adaptive large language models through cloud-edge collaboration
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c13880c8-dbdb-41ce-bce7-e66b06bc695c · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Edgeshard: Efficient LLM inference via collaborative edge computing.IEEE Internet of Things Journal, 2024
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7e64c04-0db6-4448-8b9c-9944903863d4 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2094fa64-e73e-443d-b654-c8744521eb64 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fast inference from transformers via speculative decoding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8dbfbe5d-4166-42a5-a20c-708869b7b6ef · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Accelerating large language model decoding with speculative sampling, 2023
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e31208b4-04d7-439c-be8e-37d33de933a9 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0df2741-fa12-4b04-ab32-22b12618180a · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Sequoia: Scalable, robust, and hardware-aware speculative decoding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bba04103-6a0d-43b9-b4c3-73db6ab2f825 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Lee, Deming Chen, and Tri Dao
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b204fa25-a9b6-4561-ba1f-334675d9b30d · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models EAGLE-2: Faster inference of language models with dynamic draft trees
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457a302d-9d3d-4119-ba8a-fa8b5a00785f · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Break the sequential dependency of LLM inference using lookahead decoding.arXiv preprint, 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20f5b314-7564-4fa3-a578-c7ab3c0859ad · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Reddi, and Felix Yu
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a752c88b-db56-45de-bc72-a3f4fe9e945e · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Layerskip: Enabling early exit inference and self-speculative decoding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4cab50-7dea-4036-aa84-f11339118f41 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Draft & verify: Lossless large language model acceleration via self-speculative decoding
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23008ec9-d20b-479b-b1dc-63915ef94f0b · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f66487df-1cd8-46ad-9931-9f26f325fdfb · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models QLoRA: Efficient finetuning of quantized LLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e950cda-cbcb-45d0-8480-7af5fa2c85d0 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4798fc88-828b-416f-ae75-d69c30355b48 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Improving LoRA in Privacy-preserving Federated Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b9bd033-a26f-4707-8e85-b49ebb328fd1 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models On the implicit relation between low-rank adaptation and differential privacy
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87219b76-0a36-4fad-a0be-3076474b4077 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Is split learning privacy-preserving for fine-tuning large language models?IEEE Transactions on Big Data, 2024
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a2ee5c0-71e1-41ee-ac6f-15ae965c3402 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SLDP-LoRA: A privacy-preserving split learning framework with low-rank adaptation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c36f73c1-69bf-41a8-95ca-810241f11245 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split-and-Denoise: Protect large language model inference with local differential privacy
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1483a785-638b-4235-aaee-269cc19c2a66 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Santos, et al
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc0dd2e-5e5f-4118-8583-d065b60aefcd · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models GPTQ: Accurate post-training quantization for generative pre-trained transformers
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ea7582-74b1-4a21-be37-a25d1d408574 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 143f02b4-ff8e-46cf-a862-004b59d02e04 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1bc7fc-10c5-48be-88d3-35dbd3e50c00 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models A simple and effective pruning approach for large language models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f287d4-07cc-4f10-bbab-3abaead8bc27 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Minillm: Knowledge distillation of large language models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62123e57-7a91-474e-9325-b83a21c203bc · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Efficient LLM inference on CPUs
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 715331ba-a009-48c1-a933-2b3ebe623b87 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models H2O: Heavy-hitter oracle for accurate KV cache compression
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 487eca1a-41f9-4c73-9bdd-671413163a11 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Streamingllm: Efficient streaming language models with attention sinks
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81176a82-252d-409c-ad5d-184fe5b7b48b · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Recommendation for block cipher modes of operation: Ga- lois/counter mode (gcm) and gmac
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13e26c63-4711-4bdf-bec6-4480245fafdd · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models BOLT: Privacy-preserving, accurate and efficient inference for transformers
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05b9057-df30-4519-a95a-a42138371044 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models THE-X: Privacy-preserving transformer inference with homomorphic encryption
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5369bbcc-5885-4905-9e6a-cf607f584081 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Xing, et al
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 223d5a16-0e04-49b4-8373-640f5860107b · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secureinfer: Heterogeneous TEE-GPU architecture for privacy- critical LLM tensors.arXiv preprint arXiv:2510.19979, 2025
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31bfd236-2284-4e7d-970a-f89636aeddf0 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ONNX Runtime: Cross-platform, high performance machine learning inferencing and training accelerator.https://onnxruntime.ai/
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d45801-cac2-4a28-b676-337714cc26d7 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 555f580a-0501-459e-bbc7-4be1bc3da1d2 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ML-Doctor: Holistic risk assessment of inference attacks against machine learning models
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3620fcb-bd28-4f09-92f5-16b6b7356dd0 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Privacy side channels in machine learning systems
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26177c0-e00f-4545-a5fe-dece69615396 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secure multi-party computation for machine learning: A survey.IEEE Communications Surveys & Tutorials, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2cb37e2-5ef6-48ae-8838-00b7e3d2b366 · outbound
Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.