Pith. sign in

Paper Citation Record · LEDGER

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2607.13093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.13093 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:42:19.865695Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 704996cd-c2f2-4eff-93fe-991f9522c05a · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.125500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.125500Z digest=sha256:6c8d57e67092de04af1f33901d3167aab4f34a70c2d8f939c68ca99602f64058

Observation a244b28b-a60f-4a5a-8424-34b6eba53ff3 · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.165524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.165524Z digest=sha256:42223b638c3737008ec191eafc604e70b7775ba312ef3ea824156bb64cc1d458

Observation dbf50e4c-e461-40a9-8de9-8898dd6a5caa · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Flashattention-2: Faster attention with better parallelism and work partitioning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.219713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.219713Z digest=sha256:51b9d270d6f25fc34c74a33b6fb9d0bfd28587c58c1bb6af8a9a6db62979af5c

Observation 59aafbdb-ea91-4138-abbc-039fc7e78970 · outbound

This paper cites ORCA: A distributed serving system for transformer-based generative models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ORCA: A distributed serving system for transformer-based generative models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.254849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.254849Z digest=sha256:2b2e2f77aa485c6e49f62589883e0a8114aba051fb3ca83958991a9b7f108789

Observation 9ec4aa2a-05b0-4e4d-9fa1-fbca4f3b12e6 · outbound

This paper cites DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DeepSpeed-Inference: Enabling efficient inference of transformer models at unprecedented scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.308814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.308814Z digest=sha256:32a26caf854cb924b6319c5eb17c0897fc2fc95a93560b2761c31b9a79bbe91d

Observation 7a3237f2-631f-4397-abe2-2bd07717c460 · outbound

This paper cites Neurosurgeon: Collaborative intelligence between the cloud and mobile edge.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Neurosurgeon: Collaborative intelligence between the cloud and mobile edge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.386010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.386010Z digest=sha256:a416c68cc2447784adb9fa5f80e9f1cc04a16b028ce7b2139c1038516839d0dd

Observation a2c8ebfc-c67a-4f6c-8050-c791406b6084 · outbound

This paper cites Split computing and early exiting for deep learning applications: Survey and research challenges.ACM Computing Surveys, 2022.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split computing and early exiting for deep learning applications: Survey and research challenges.ACM Computing Surveys, 2022

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.418410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.418410Z digest=sha256:41566e00950f4d9e81e386d0866b075f6ec8dc2508607bc8007e07e9feeef8d4

Observation be4ee5b6-3c18-4bc6-ace7-ce68744ba00a · outbound

This paper cites DNN surgery: Accelerating DNN inference on the edge through layer partitioning.IEEE Transactions on Parallel and Distributed Systems, 2023.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DNN surgery: Accelerating DNN inference on the edge through layer partitioning.IEEE Transactions on Parallel and Distributed Systems, 2023

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.469966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.469966Z digest=sha256:41ac7974d68188a7e0670ed4ef1afd9cbdac99aa687be8d5c2b7f2437c38fdef

Observation 5a7f8188-0c65-4f6f-be70-bde979f3536d · outbound

This paper cites Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Pipeedge: Pipeline parallelism for large-scale model inference on heterogeneous edge devices

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.527834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.527834Z digest=sha256:5ac47887748ef04729fe78abb158af669c84b8b273661afe72da21df586e9be0

Observation a3e7d1d7-cc21-4428-a652-4b0f664a237c · outbound

This paper cites Hybrid SLM and LLM for edge-cloud collaborative inference.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hybrid SLM and LLM for edge-cloud collaborative inference

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.597558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.597558Z digest=sha256:6a0511a1147fbdb8b71c4f8adec00b43d49bbb873d80974253ccdd19b3802c58

Observation 33614f3b-2045-4f8a-b6b4-4898256a8cef · outbound

This paper cites Jupiter: Fast and resource-efficient collaborative inference of generative LLMs on edge devices.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Jupiter: Fast and resource-efficient collaborative inference of generative LLMs on edge devices

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.654481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.654481Z digest=sha256:bcc61fcade1e7b4b118809fdc91a539b8521f6ba5286ca31205f077c3507dfa4

Observation 50603c84-6301-4705-b54b-b447c97d11af · outbound

This paper cites DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models DSSD: Efficient Edge-Device LLM Deployment and Collaborative Inference via Distributed Split Speculative Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.700801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.700801Z digest=sha256:5d6c5d772d0ef86bf10bd038ce739064b860ced11165745c7de1e80a53aa5cbd

Observation cbd50d06-12a4-4179-8465-1e9474670daa · outbound

This paper cites CE-CoLLM: Efficient and adaptive large language models through cloud-edge collaboration.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models CE-CoLLM: Efficient and adaptive large language models through cloud-edge collaboration

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.760422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.760422Z digest=sha256:54352c7cc5549d1179b6ba324c64fe6f750497b404aaa74425e66a9638b2b51b

Observation c13880c8-dbdb-41ce-bce7-e66b06bc695c · outbound

This paper cites Edgeshard: Efficient LLM inference via collaborative edge computing.IEEE Internet of Things Journal, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Edgeshard: Efficient LLM inference via collaborative edge computing.IEEE Internet of Things Journal, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.840945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.840945Z digest=sha256:7c6371f5e74d2148803ebd4579aaa34f63f7e69c71fcbde15aa4795beee57037

Observation a7e64c04-0db6-4448-8b9c-9944903863d4 · outbound

This paper cites Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Collaborative Inference and Learning between Edge SLMs and Cloud LLMs: A Survey of Algorithms, Execution, and Open Challenges

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.915097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.915097Z digest=sha256:5a3ccf55d609eaa4fd082994db61ab99305807aaa2ad9cf95168134869511683

Observation 2094fa64-e73e-443d-b654-c8744521eb64 · outbound

This paper cites Fast inference from transformers via speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Fast inference from transformers via speculative decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:15.997029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:15.997029Z digest=sha256:3fb94bef219067bea53b881382a0b06b7b6712c8fa7842d03fc91c4961e192ba

Observation 8dbfbe5d-4166-42a5-a20c-708869b7b6ef · outbound

This paper cites Accelerating large language model decoding with speculative sampling, 2023.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Accelerating large language model decoding with speculative sampling, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.090265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.090265Z digest=sha256:57f0465b79bb2a0ee964f2a2233e7859009d41861f206aab8035cc3b9d3bfedf

Observation e31208b4-04d7-439c-be8e-37d33de933a9 · outbound

This paper cites Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Specinfer: Accelerating generative large language model serving with tree-based speculative inference and verification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.143946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.143946Z digest=sha256:4156ecafb211537107d8123f5d240c45646f7407643b9306e3f1c1f0a1c126c9

Observation b0df2741-fa12-4b04-ab32-22b12618180a · outbound

This paper cites Sequoia: Scalable, robust, and hardware-aware speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Sequoia: Scalable, robust, and hardware-aware speculative decoding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.237006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.237006Z digest=sha256:54e5217dc5a6501db5e77aa92becda124e277d26a7d722324ef316dcea0e37aa

Observation bba04103-6a0d-43b9-b4c3-73db6ab2f825 · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Lee, Deming Chen, and Tri Dao

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.304835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.304835Z digest=sha256:989942ed191678e26893a61e5a12e96e034c6f258ebedb01dbb90d52e56dfe05

Observation b204fa25-a9b6-4561-ba1f-334675d9b30d · outbound

This paper cites EAGLE-2: Faster inference of language models with dynamic draft trees.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models EAGLE-2: Faster inference of language models with dynamic draft trees

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.363212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.363212Z digest=sha256:21c41ecd16696660f67e6c54010573b3a3245f429ff07934359ed02a9299d798

Observation 457a302d-9d3d-4119-ba8a-fa8b5a00785f · outbound

This paper cites Break the sequential dependency of LLM inference using lookahead decoding.arXiv preprint, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Break the sequential dependency of LLM inference using lookahead decoding.arXiv preprint, 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.391598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.391598Z digest=sha256:80f51d835b60556e587e4be1e444c983ecc41e42639383a9f6bf2ab1fb8eea8b

Observation 20f5b314-7564-4fa3-a578-c7ab3c0859ad · outbound

This paper cites Reddi, and Felix Yu.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Reddi, and Felix Yu

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.475568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.475568Z digest=sha256:5f47b7f4c2318f0cf51feaebc2cb4de2d0888ff335d181b9f26cffdf767c56b1

Observation a752c88b-db56-45de-bc72-a3f4fe9e945e · outbound

This paper cites Layerskip: Enabling early exit inference and self-speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Layerskip: Enabling early exit inference and self-speculative decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.542612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.542612Z digest=sha256:083435e6f3dff12f00c34dd6cfea3148da54c64920fb0ef15a385a6852836e53

Observation 7c4cab50-7dea-4036-aa84-f11339118f41 · outbound

This paper cites Draft & verify: Lossless large language model acceleration via self-speculative decoding.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Draft & verify: Lossless large language model acceleration via self-speculative decoding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.615010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.615010Z digest=sha256:c90b2c76e01d8a05047daac93fc314984794201b7a07f2964bb5c01ae0370f9d

Observation 23008ec9-d20b-479b-b1dc-63915ef94f0b · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.702536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.702536Z digest=sha256:c4d1c91edd7c5e2ad69620bbe632a3db969d9f34ecf956b0898d195adf396ab6

Observation f66487df-1cd8-46ad-9931-9f26f325fdfb · outbound

This paper cites QLoRA: Efficient finetuning of quantized LLMs.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models QLoRA: Efficient finetuning of quantized LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.739903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.739903Z digest=sha256:c940827c3d6229fbcb775ffddb3f1237a9fd1e286b000b6f624c1c17407e6ce6

Observation 8e950cda-cbcb-45d0-8480-7af5fa2c85d0 · outbound

This paper cites SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SplitLoRA: A Split Parameter-Efficient Fine-Tuning Framework for Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:16.876672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:16.876672Z digest=sha256:4963bf715e9b4e62d1ae3ede9ab64555847aeefbdd8ec1c3f8c06f79f4f91c4f

Observation 4798fc88-828b-416f-ae75-d69c30355b48 · outbound

This paper cites Improving LoRA in Privacy-preserving Federated Learning.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Improving LoRA in Privacy-preserving Federated Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.098363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.098363Z digest=sha256:e98a3303302958338c0092ec128933ddbe1471d9a9abedc83486eec310af3564

Observation 9b9bd033-a26f-4707-8e85-b49ebb328fd1 · outbound

This paper cites On the implicit relation between low-rank adaptation and differential privacy.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models On the implicit relation between low-rank adaptation and differential privacy

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.249674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.249674Z digest=sha256:0efff47486862b92da1ae80e40b5908b2feab47811b01d8bf98d6e94f99931ac

Observation 87219b76-0a36-4fad-a0be-3076474b4077 · outbound

This paper cites Is split learning privacy-preserving for fine-tuning large language models?IEEE Transactions on Big Data, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Is split learning privacy-preserving for fine-tuning large language models?IEEE Transactions on Big Data, 2024

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.384853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.384853Z digest=sha256:a5b993689cbe9943d9490d3f7bd6e2d567c347655edde270cfd63bf3b03a77a1

Observation 8a2ee5c0-71e1-41ee-ac6f-15ae965c3402 · outbound

This paper cites SLDP-LoRA: A privacy-preserving split learning framework with low-rank adaptation.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models SLDP-LoRA: A privacy-preserving split learning framework with low-rank adaptation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.521863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.521863Z digest=sha256:eb3645cbb0c36638f2b1e569050fd9bd3674e50f176423850d1d57f2762a1c92

Observation c36f73c1-69bf-41a8-95ca-810241f11245 · outbound

This paper cites Split-and-Denoise: Protect large language model inference with local differential privacy.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Split-and-Denoise: Protect large language model inference with local differential privacy

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.678536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.678536Z digest=sha256:38d950f6a8c54d5b048531ea39b3bb8625d5e0561734c854031b879adb97dfa8

Observation 1483a785-638b-4235-aaee-269cc19c2a66 · outbound

This paper cites Santos, et al.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Santos, et al

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:17.851965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:17.851965Z digest=sha256:5ae1b28e4e2b119316d6849c4a5aba8d46ab74df2b146481c139df334a60da14

Observation 2dc0dd2e-5e5f-4118-8583-d065b60aefcd · outbound

This paper cites GPTQ: Accurate post-training quantization for generative pre-trained transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models GPTQ: Accurate post-training quantization for generative pre-trained transformers

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.076697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.076697Z digest=sha256:e1763b9ccf75e0d319dbc626441bb26188bc7db9d6420f46708800a785e25787

Observation 67ea7582-74b1-4a21-be37-a25d1d408574 · outbound

This paper cites AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models AWQ: Activation-aware weight quantization for on-device LLM compression and acceleration

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.194038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.194038Z digest=sha256:8c1fb515361ff2a7a72f235fa7deaf1fad312dd1128362a6361ee41c8b039b14

Observation 143f02b4-ff8e-46cf-a862-004b59d02e04 · outbound

This paper cites Smoothquant: Accurate and efficient post-training quantization for large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Smoothquant: Accurate and efficient post-training quantization for large language models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.257172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.257172Z digest=sha256:5c3673cd012d2b36f971c337dc755ce02f706b8c91a29e8a9cec550d325074cf

Observation 7c1bc7fc-10c5-48be-88d3-35dbd3e50c00 · outbound

This paper cites A simple and effective pruning approach for large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models A simple and effective pruning approach for large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.321759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.321759Z digest=sha256:db2651e1e40ca3e86fdfda67361902564120c22f15ceb40b055629a1e4b90f90

Observation 58f287d4-07cc-4f10-bbab-3abaead8bc27 · outbound

This paper cites Minillm: Knowledge distillation of large language models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Minillm: Knowledge distillation of large language models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.401946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.401946Z digest=sha256:36b212358ba8c58a02cf6c607fa087000f28b9c731434b263470e9f535047066

Observation 62123e57-7a91-474e-9325-b83a21c203bc · outbound

This paper cites Efficient LLM inference on CPUs.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Efficient LLM inference on CPUs

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.479747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.479747Z digest=sha256:764342194df697cce642a388d61f5eb541cd53b34f2c2a8fb8d00c2bb27093ed

Observation 715331ba-a009-48c1-a933-2b3ebe623b87 · outbound

This paper cites H2O: Heavy-hitter oracle for accurate KV cache compression.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models H2O: Heavy-hitter oracle for accurate KV cache compression

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.566986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.566986Z digest=sha256:a7341972d4a2fd25bca64a0db65cb04a7204588ed10a9bcdabd0bb11a75751b2

Observation 487eca1a-41f9-4c73-9bdd-671413163a11 · outbound

This paper cites Streamingllm: Efficient streaming language models with attention sinks.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Streamingllm: Efficient streaming language models with attention sinks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.679844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.679844Z digest=sha256:83e01c8afb45114ac8737741c1421d2ad873f3471dbe731a889c7a091394c8a0

Observation 81176a82-252d-409c-ad5d-184fe5b7b48b · outbound

This paper cites Recommendation for block cipher modes of operation: Ga- lois/counter mode (gcm) and gmac.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Recommendation for block cipher modes of operation: Ga- lois/counter mode (gcm) and gmac

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.791543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.791543Z digest=sha256:b40df9a31a1c6415366794f8aab78e2d157c50d27cbd45a853e1a4d42e18691e

Observation 13e26c63-4711-4bdf-bec6-4480245fafdd · outbound

This paper cites BOLT: Privacy-preserving, accurate and efficient inference for transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models BOLT: Privacy-preserving, accurate and efficient inference for transformers

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.867792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.867792Z digest=sha256:078ce8758381996c3823301651afef8ed93418352f3f95ef2f0b5269a5595e89

Observation e05b9057-df30-4519-a95a-a42138371044 · outbound

This paper cites THE-X: Privacy-preserving transformer inference with homomorphic encryption.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models THE-X: Privacy-preserving transformer inference with homomorphic encryption

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:18.969333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:18.969333Z digest=sha256:8fbb2569fc46fcc48e1e6276f26a8ebc629f03f79f688897d63c96e79590bbb3

Observation 5369bbcc-5885-4905-9e6a-cf607f584081 · outbound

This paper cites Xing, et al.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Xing, et al

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.083020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.083020Z digest=sha256:04d3967dec61b8dadd107ac3ca86e0d6d264b9acf3633797e653c789be7ae824

Observation 223d5a16-0e04-49b4-8373-640f5860107b · outbound

This paper cites Secureinfer: Heterogeneous TEE-GPU architecture for privacy- critical LLM tensors.arXiv preprint arXiv:2510.19979, 2025.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secureinfer: Heterogeneous TEE-GPU architecture for privacy- critical LLM tensors.arXiv preprint arXiv:2510.19979, 2025

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.164171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.164171Z digest=sha256:ec6e6ddc0d515622217848c391d3a24c68a72a15c40d5bcf88d265f1a9ed352c

Observation 31bfd236-2284-4e7d-970a-f89636aeddf0 · outbound

This paper cites ONNX Runtime: Cross-platform, high performance machine learning inferencing and training accelerator.https://onnxruntime.ai/.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ONNX Runtime: Cross-platform, high performance machine learning inferencing and training accelerator.https://onnxruntime.ai/

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.238337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.238337Z digest=sha256:98967d0bbec0a218b127170253a605df86112bc243b4d8d3524d2d5811299356

Observation 02d45801-cac2-4a28-b676-337714cc26d7 · outbound

This paper cites Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Text Revealer: Private Text Reconstruction via Model Inversion Attacks against Transformers

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.309412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.309412Z digest=sha256:0328043c6678ee0b6eb503e3213f57eb03fc0bb580c040166b04070097e20d70

Observation 555f580a-0501-459e-bbc7-4be1bc3da1d2 · outbound

This paper cites ML-Doctor: Holistic risk assessment of inference attacks against machine learning models.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models ML-Doctor: Holistic risk assessment of inference attacks against machine learning models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.410428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.410428Z digest=sha256:08d6fa3460da73d4c910dbc477aea2f9eb03c7bc64fa5fe9479c29d4b6716d45

Observation b3620fcb-bd28-4f09-92f5-16b6b7356dd0 · outbound

This paper cites Privacy side channels in machine learning systems.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Privacy side channels in machine learning systems

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.565403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.565403Z digest=sha256:7f6c2909cc452f99c6e12afc118f842858190a084056ac9f70df9744dbca6a45

Observation f26177c0-e00f-4545-a5fe-dece69615396 · outbound

This paper cites Secure multi-party computation for machine learning: A survey.IEEE Communications Surveys & Tutorials, 2024.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Secure multi-party computation for machine learning: A survey.IEEE Communications Surveys & Tutorials, 2024

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.737748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.737748Z digest=sha256:23f8a53acae40c64c51e89a0c77b5a64f7e48aa4d6958fe4dec9ce413b175e1c

Observation a2cb37e2-5ef6-48ae-8838-00b7e3d2b366 · outbound

This paper cites an unresolved cited work.

Efficient and Privacy Aware Edge Cloud Collaborative Inference for Large Language Models Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T06:42:19.865695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:42:19.865695Z digest=sha256:ec9fe303df6d48d5f27e2fd868b128b1d58ceb77cd334fe115112d8de1249a61

Pith citing papers

No inbound Pith citation observations are available.