Pith. sign in

Paper Citation Record · LEDGER

LLM Inference Unveiled: Survey and Roofline Model Insights

As of 4 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 35 inbound Pith citation observations for arXiv:2402.16363.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.16363 v6

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 35 of 35 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T16:35:33.896184Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T01:57:51.001945Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 90806b5a-d669-4a41-ae24-75184098e22d · inbound

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models cites this paper.

ASVD: Activation-aware Singular Value Decomposition for Compressing Large Language Models LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:49:33.794422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T13:49:33.747672Z digest=sha256:7a7a09a8d49b8f8e0945405e0bf60233816a0fa50e8be2b44d4281c17547474e

Observation 456ea09f-255f-4c50-a111-d7d38c867900 · inbound

A Survey on Efficient Inference for Large Language Models cites this paper.

A Survey on Efficient Inference for Large Language Models LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:39:33.159051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-15T02:39:33.007894Z digest=sha256:6b075e7b41de26a7aedcbf7c1fcaeeea500f2852bbf233ac35e8986b8802bb32

Observation a7a73652-6c22-41ac-92fe-6bfd2dedddae · inbound

HybridFlow: A Flexible and Efficient RLHF Framework cites this paper.

HybridFlow: A Flexible and Efficient RLHF Framework LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:53:38.959988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-11T07:53:38.715353Z digest=sha256:07c6bfcd6982436439dec1e1bed1d533ad253983e7419e8160f1aee9a3cda0fc

Observation 70f92a0a-1dbc-4fd9-840c-8d0f1db5a590 · inbound

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation cites this paper.

LogQuant: Log-Distributed 2-Bit Quantization of KV Cache with Superior Accuracy Preservation LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-22T22:12:11.621335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-22T22:07:37.906916Z digest=sha256:d5db844ec8a7e6c4e59d1fa30c899641ad72be762893807f22cc26d9f1320abb

Observation 56688505-a8b2-4bcc-9b9b-e778f36b9c4e · inbound

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference cites this paper.

HPIM: Heterogeneous Processing-In-Memory-based Accelerator for Large Language Models Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T16:35:33.896184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T16:35:33.896184Z digest=sha256:d0b49017d887324360c41e2b65876139de4d85f6a8413b81bb932a428d78cce3

Observation 159f7077-0ace-467e-af00-344d9027a091 · inbound

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads cites this paper.

Analytical Provisioning for Attention-FFN Disaggregated LLM Serving under Stochastic Workloads LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:57:43.030400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T09:53:13.057587Z digest=sha256:24d53178fd2bbb17ebb68b68caff42317969c2ee319a2e55e59d4dca124153ab

Observation 0d18adca-0792-44d3-890d-59c722e37061 · inbound

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference cites this paper.

SweetSpot: An Analytical Model for Predicting Energy Efficiency of LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:07:29.606715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-16T07:05:56.380048Z digest=sha256:a962a20385960595b1828f937128ce5409f6e2c06a5bb3b302be7c08ef401b8f

Observation 147cae87-5bfc-4a14-beb9-9ecd4747d509 · inbound

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs cites this paper.

ENEC: A Lossless AI Model Compression Method Enabling Fast Inference on Ascend NPUs LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:58:03.122102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-14T21:56:38.684104Z digest=sha256:ec9069646a3512fc25ce34bdca05ed3ab6976ca753e48ccf48e92d20449605be

Observation f5ff9e9a-1a6d-4c80-ac42-0fe94c2f005e · inbound

Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling cites this paper.

Cactus: Accelerating Auto-Regressive Decoding with Constrained Acceptance Speculative Sampling LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:38:02.718186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-13T17:36:06.003661Z digest=sha256:2c91e3d2a37df52319441caf6e726facc230c0a0b3f9f99ee665078f1ac18683

Observation a65b5d76-97c6-4b4b-a227-02f450f67153 · inbound

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics cites this paper.

Autopoiesis: A Self-Evolving System Paradigm for LLM Serving Under Runtime Dynamics LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:19.616843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T17:30:40.021376Z digest=sha256:5dccfd312ed49bdd0cf46fd87eeb11621d006cd815060348bb8e267ce0d109a5

Observation 045be78f-94a9-4672-9169-9c9e6102a303 · inbound

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding cites this paper.

Why and When Visual Token Pruning Fails? A Study on Relevant Visual Information Shift in MLLMs Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.463875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-10T14:50:37.022338Z digest=sha256:06307d059f13fa9f5f8f12553f5dce01579c495844d0e3bc25003d57040518af

Observation 056bba68-557b-4834-9c14-b507ffce08df · inbound

A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks cites this paper.

A Task Decomposition and Planning Framework for Efficient LLM Inference in AI-Enabled WiFi-Offload Networks LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:12.821198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:11:26.862363Z digest=sha256:74e842d68f2c491c069a97ca116ab39a598ad426f377a5ab9cf8a3ab729a612f

Observation 8ef061b4-f73b-4e58-80d3-c1a271d5a89c · inbound

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding cites this paper.

Making Every Verified Token Count: Adaptive Verification for MoE Speculative Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:31:05.348415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-09T19:53:40.791974Z digest=sha256:53213d285f5e2d9bc8777ad74c32c0af646da5b00528e33180152904b2f24ad6

Observation 06e377ad-dcda-403a-bccc-88435008d2aa · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:26.967126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-09T18:43:06.803528Z digest=sha256:1d04f7a2a017acb3c23a3a632b3c79b2ec19e7e7958cc0fb41b62e84df027aa0

Observation e530f9f2-06aa-4692-912e-9a96b177b4c0 · inbound

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis cites this paper.

Sim-FA: A GPGPU Simulator Framework for Fine-Grained Asynchronous Pipeline Analysis LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T15:09:44.098075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:09:44.098075Z digest=sha256:20e6ab2ca98a7e09e64ab46515ca05d5683c3eed6d0115ebf17a181f23dffd59

Observation e9c5676f-0207-4e02-bdfb-cf1ed7537661 · inbound

Gated Subspace Inference for Transformer Acceleration cites this paper.

Gated Subspace Inference for Transformer Acceleration LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:31.916138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T19:26:08.510348Z digest=sha256:32810a71ac158de0ab904eccba01b5639047f0661066894d000de906b4588b77

Observation 42e9d4d3-7ed4-41d9-a55b-429b60591205 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:30:44.191274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T18:23:14.935801Z digest=sha256:f1c6bee80d52d2e4474e6ae31b9b0399218bd59685f6ed1019acba8116db0a70

Observation b24b8616-dbc5-43d5-a0c1-d1b4ed63a515 · inbound

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization cites this paper.

OSAQ: Outlier Self-Absorption for Accurate Low-bit LLM Quantization LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:18.177694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-12T02:59:00.997742Z digest=sha256:e8c229af01271be0ab1c799c468fb5a687085184e751b48186685b23b2f4b49f

Observation 8246754b-f481-416b-9f79-16f07c5cff9c · inbound

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference cites this paper.

KVDrive: A Holistic Multi-Tier KV Cache Management System for Long-Context LLM Inference LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:18:13.959019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-20T11:13:46.098095Z digest=sha256:ac24217c5a742d11f2647340cadf6e153c69b7db51a3e5205d73b3e8b60bf79d

Observation 1bd1c73a-6b8a-45a0-a48f-b7e1d9ba5b0c · inbound

On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection cites this paper.

On-Device Generative AI for GDPR-Compliant Visual Monitoring: Natural Language Alerts from Local Object Detection LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.308427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T07:51:13.367200Z digest=sha256:01962e0feb5631e1681297038ec5e8f1050ad7ae3bc72b3a0835e996e5e88219

Observation 275b708e-8721-476b-a29a-1fd21f16047b · inbound

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture cites this paper.

Model-Native Computing Architecture: Envisioning Future System Architecture Through the Lens of Computer Architecture LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:36:09.263898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T22:12:13.114405Z digest=sha256:965ca94752c41209a47a19ff64a69f136fe86c3c75129047117b8f12f79c0e25

Observation 5bb01164-4430-424f-8618-806e1500196c · inbound

EinSort: Sorting is All We Need for Tensorizing LLM cites this paper.

EinSort: Sorting is All We Need for Tensorizing LLM LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:57:26.711487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-06-27T18:31:01.804061Z digest=sha256:083554a52fe3708e3959096ea061b74b59936f0779129e801ed83a623dbdffd7

Observation edaa94ed-94a3-4f1c-85da-db9cf5bf72a7 · inbound

Operator Fusion for LLM Inference on the Tensix Architecture cites this paper.

Operator Fusion for LLM Inference on the Tensix Architecture LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:16:44.969993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T07:01:15.392433Z digest=sha256:cf5ee0c262ac6da4a13112d10a83eb0051fff600d7dd440ff66c3e933ce2a359

Observation d58a23c3-f7e6-4a53-997d-a141d21b67c9 · inbound

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts cites this paper.

EfficientRollout: System-Aware Self-Speculative Decoding for RL Rollouts LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:59:07.177818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T21:32:41.428871Z digest=sha256:ec76813f526fe496cf3622a040a85dd0b40572d994a4e29ef678b7ec33314e4c

Observation 0c1ecd48-3c85-496f-86e8-8da5a2c8fcfb · inbound

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding cites this paper.

Dustin: Draft-Augmented Sparse Verification for Efficient Long-Context Generation with Speculative Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:49:57.985869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-26T00:11:58.636939Z digest=sha256:89d032f9c594f8ac3e0650943bf21fdb521c8764b25ba8df6e7735796441d4a4

Observation 71735d02-751c-4462-b4a0-537cf837c9c6 · inbound

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression cites this paper.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:29:22.036003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-04T01:23:46.175022Z digest=sha256:8e7065873c5e4fd88548c25b77f8c4c1d28ee2b7870e8876bac7b8d2ec48bc73

Observation 335dbaa4-29f2-468a-a406-8b2e36a447af · inbound

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression cites this paper.

KARA: Efficient Reasoning LLM Serving via Sliding-Window KV Cache Compression LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T17:54:40.241144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:54:40.241144Z digest=sha256:96f546b5de85e5a91b50221be8d527682282137b820a1dfdef4a7664b6b6d9a1

Observation e4ae8fa4-9ab3-4f41-9429-8aca7ece18b7 · inbound

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems cites this paper.

MLSYSIM: First-Principles Infrastructure Modeling for Machine Learning Systems LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 136

Resolution
unresolved
no resolver link, observed 2026-07-12T11:05:56.233115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T11:05:56.233115Z digest=sha256:149a4b9dfa7918df0df9f7f7538cfe6098a00f851a963373360c2213de97e21e

Observation 3df85082-7be3-42d5-bf95-d3b761ddf1b2 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-08T22:45:40.136167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-08T22:38:12.637901Z digest=sha256:432375dd0d2b216cc1710c21e0f3e134232469440b116584c630782e109188f8

Observation a0405de9-45c6-4168-a446-b473c3262421 · inbound

Think Before You Grid-Search: Floor-First Triage for LLM Serving cites this paper.

Think Before You Grid-Search: Floor-First Triage for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-11T01:57:51.036518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-11T01:55:09.658053Z digest=sha256:b0f619d3f35f12d0ec7fb8c1f78645e4611d658e241952ce80d41035799f015e

Observation 8ecd1f7a-97e0-4398-8016-4c8f0275c8f9 · inbound

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing cites this paper.

A Blueprint for Equilibrium-Based Differentiable Continuous-Variable Thermodynamic Computing LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:12.343498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:12.343498Z digest=sha256:1113d735e315177bc9cf04be6412764f0a17f19c9f850917818736847ff5b3d9

Observation 9168e695-d063-4ad4-9634-4fa85f8a9393 · inbound

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding cites this paper.

HiKV: Hierarchical Importance-Aware KV Cache with Hardware Acceleration for LLM Decoding LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T04:58:46.181845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:58:46.181845Z digest=sha256:c35bbd9cac99558c67cc63c7293f5b1042f5b040a18b120b3cf6166357310264

Observation 7af833aa-6c1b-444c-9024-198ac481d754 · inbound

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters cites this paper.

TileSight: A First-Principles Tile-Centric Analytical GPU Performance Model from Cores to Clusters LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-01T04:49:48.043261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T04:49:48.043261Z digest=sha256:92b623c7e5299c09a89b2844d6d77cae9ace35931d4732c39ba9394c478e29c9

Observation 9d2d8d7c-48a5-4710-a506-4a869d174421 · inbound

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement cites this paper.

NELSSA: A GPU-PNM Heterogeneous System for Mixed-Length LLM Serving via Length-based Request Placement LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T12:00:24.412238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:00:24.412238Z digest=sha256:e8c9ab118f19e5bfd8c7ef911c373caa0132dae7a9aab68ac88fe70057e5e9c3

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · inbound

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving cites this paper.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:dae31b75d359011c998d27809790238df0ab30f84b0761fadc842338c50a3e25