Pith. sign in

Paper Citation Record · LEDGER

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

As of 15 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.29575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29575 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1283f76c-0bbf-4207-80e6-721c88cf4e48 · outbound

This paper cites The rapid adoption of generative ai,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The rapid adoption of generative ai,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.621032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.621032Z digest=sha256:29b8288f362646eada878244886f1bcbfab272a2aad79f6282d2256b62e567a5

Observation 71750f8e-e65d-4dfc-8fc7-5c8e831faba4 · outbound

This paper cites The adoption of chatgpt,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The adoption of chatgpt,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.687853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.687853Z digest=sha256:ebe1baa3801e036214170646a3271639b854c88e7c87c51004b26c6212a69722

Observation 5e3aec83-1a24-4388-adde-a9bb58b1fdf0 · outbound

This paper cites Quantifying large language model usage in scientific papers,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Quantifying large language model usage in scientific papers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.737457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.737457Z digest=sha256:1007a8f02ac2c2c739845d0eba06ed040f41a6b53a5979c2a56d497e511f054b

Observation d7a31b72-eb9d-4ed6-8138-bb0287ef2850 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llumnix: Dynamic scheduling for large language model serving,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.844613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.844613Z digest=sha256:3a89f9ac358bd40e1f5e3ae1fdb2bf3b124e7e7fa37849645937fc387c00a80c

Observation a627426c-a4f4-4206-a394-04d805de12ae · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.935319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.935319Z digest=sha256:c66aef74ff5fd9e284404ffac61f383a3d27236c3b0a7ccf12631abe2918c56f

Observation dc0fee93-9e4d-46c1-8b28-cfa5f828b8ae · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.032744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.032744Z digest=sha256:b23e4c34f799c3a385e3f53e284502be1dd09cabff8815fbf327752d981ee5a4

Observation 3d7e53ee-93d8-47cb-a734-76e6e4e5c880 · outbound

This paper cites Efficiently scaling transformer inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficiently scaling transformer inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.082197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.082197Z digest=sha256:b817c586e9f4421f6251944eefc882dfcf0570a44b3d5ab6fca77fcce0c4d248

Observation dbd5fb7e-2239-4859-89da-77361ba9beb1 · outbound

This paper cites Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.167267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.167267Z digest=sha256:fc4921eb349332fdae77d1e8b867ea49468800f6524c8cc67238b8bef7915bc8

Observation 8401ac48-13a1-42e6-96f7-d75a09241a32 · outbound

This paper cites Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.228945Z digest=sha256:2ec517ea78b762f020ccaba2c1ec792cdd819d3b1329e70bdf97c491f29d97f1

Observation e9be6525-2ed8-4deb-baa1-20852599966e · outbound

This paper cites Llmvisor: A real-time latency attribution model for multi-tenant llm serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmvisor: A real-time latency attribution model for multi-tenant llm serving,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.312448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.312448Z digest=sha256:074884ba3fc1e46818ed1eca00a0ec0b0e93260cebbe20951e5fe98fefa1dcf4

Observation 169dd382-9cd3-48bf-8c43-d32d50983289 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Predicting llm inference latency: A roofline-driven ml method,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.386461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.386461Z digest=sha256:befffd125eb6c8d7080b35baf7fc2e1867ce7cfc9054fab8dc51ce2c4b7d3b77

Observation faa6fa12-8937-465d-87bf-2c10a68eb48a · outbound

This paper cites Language mod- els are few-shot learners,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Language mod- els are few-shot learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.529800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.529800Z digest=sha256:cdb78c1203712cb89fceb58a69a320ac9fe95d756c9fd4aa98c4e67f48b091af

Observation 06875560-041e-47bd-bca5-4e0a529eaf44 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.616647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.616647Z digest=sha256:4483efd8b4f5cc153531ed7e1fb47860226ef7a9b24846836da62fac88c43a5e

Observation ab551368-fb6b-46ed-b8d7-789c77352651 · outbound

This paper cites Attention is all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.703801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.703801Z digest=sha256:e8ed3abaff4f8e62b28b26f8adbcc36e5bf293473f17b1e3326350cb623a3d2d

Observation de2897af-d4a1-4f69-8e86-b899daadc481 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Orca: A distributed serving system for transformer-based generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.769107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.769107Z digest=sha256:b7ea124c8a371b3c5cb037e24f8dd2e08eb2ed12722976019c5067986d4b52cf

Observation 56f5eeae-c9e5-433a-b6c4-22b4f42723c9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient memory management for large language model serving with pagedattention,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.843887Z digest=sha256:1c6b24645a3876eaea7bac0f6c93858d80888f5bdd0369b15e8e65ca75b75df9

Observation 3aa6c78b-0554-4beb-a73e-df1f537fc2e2 · outbound

This paper cites TensorRT-LLM,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving TensorRT-LLM,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.928323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.928323Z digest=sha256:e3e347e37e29a123a7d657174b4406e2b1412cf7dcc10b59d0b34cf6b11ccded

Observation 94854dc3-b780-4314-b2a6-ef0ecc871228 · outbound

This paper cites DeepSpeed-MII,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving DeepSpeed-MII,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.030125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.030125Z digest=sha256:928b11fc863cba9b5031b6930548fd02327dbd4d79504d542b032cfbf2c34ba2

Observation f142fcee-c690-41b7-a722-d974ca81131e · outbound

This paper cites Slora: Scalable serving of thousands of lora adapters,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Slora: Scalable serving of thousands of lora adapters,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.117570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.117570Z digest=sha256:4d6f3d237a9809db40bfcde0fcd10ec99cf6e38686be42b1b8129b67a90c6d15

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:9dd591fdae0fe675c81c7cb8fb4a5165421c948dba2b795c17ed0048c745c80f

Observation 46ded418-0c69-4282-b2fd-5ca8d11f8edc · outbound

This paper cites Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.305385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.305385Z digest=sha256:0bac58a775446006aa691c3969ff6348a86531e3c61881c040546ae95bc86d15

Observation e556d816-ff08-40d6-835e-f20286af23c8 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast inference from transform- ers via speculative decoding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.428587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.428587Z digest=sha256:f30c9ad6919a2feccf75d6f5383600306cb8940ca3e16589c7f6905b3a1ecd17

Observation 11947e0b-a77a-40a4-9b7f-51eeb4fa4fc2 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.550502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.550502Z digest=sha256:e03cf1775b5675baceae97f3d3a3aa4e80e0465fb31f3463d3235b1eeb0aa01c

Observation f0c8d29b-ce4a-4423-b2d6-fc9afdaae304 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast Transformer Decoding: One Write-Head is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.668526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.668526Z digest=sha256:adc39dbb7dd72dde3b14c35230fe7ee8b202710ce6bcb3e106161d9c344bdebe

Observation 4fff4c06-15c0-4d6c-82a7-e8a0c5e61c97 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.767380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.767380Z digest=sha256:f658209d5f7dc2803878e39cbc4cd28c16f960268d2e2dcbb1543f44a6357c8f

Observation 8a9d34a5-5e25-41b3-b4f1-f9dde91eb23f · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.847735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.847735Z digest=sha256:e67a99f4b85b3de105ff929fed48daded25fdffa435c12acfcfd1d1ddf3d2ec0

Observation 6f35459a-2052-4824-bc59-7fc008a098fc · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.915456Z digest=sha256:cb41f48fcc0e9fc79136b8e208801bd81e1cf7af495199716aa82c4546167db3

Observation ec390e4f-1a26-45ca-b8ef-c96e3928b659 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Vidur: A large-scale simulation framework for llm inference,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.024548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.024548Z digest=sha256:9ae5bd7d9340615ef075b5cb7206f288563061a18ee55c32b36cd8c176741aea

Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.136193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.136193Z digest=sha256:18766342e85705675fed3297a2315f2096ee5512aa8db976f56762f5d3c3c8f0

Observation e3a09dfd-6bd0-4c5c-b9c0-dad1f0f4702a · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.209163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.209163Z digest=sha256:1397995f1adc372f7ebcce3a55419c6c9a1f5dfe2e4afc3c9075fe2c49fb2096

Observation 8d956b91-494c-4617-83d1-4b357831ad48 · outbound

This paper cites Amali: An analytical model for accurately modeling llm inference on modern gpus,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Amali: An analytical model for accurately modeling llm inference on modern gpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.321060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.321060Z digest=sha256:c719c258514233667fa37a5c7d02b5593fac56731e9634fc874a99b0b719e837

Observation 46db8890-3e1f-4963-9f18-6cc14c0f8fef · outbound

This paper cites Fairness in serving large language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fairness in serving large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.393342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.393342Z digest=sha256:54caa24b29bae929e50746f4dbac31a53f24685780b2cc65f2b0bb59d92c8b7d

Observation ef1c6471-f486-4f1d-a3b2-dc42413c0c53 · outbound

This paper cites Clean sharegpt dataset,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Clean sharegpt dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.458412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.458412Z digest=sha256:28345227bfb93d168001bdf58e3bc16f9c69f3900f0d90c6123fc29bb17196e7

Observation 6b314849-88b5-42a6-910a-381b7b4b23f0 · outbound

This paper cites Mistral 7B.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mistral 7B

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.533081Z digest=sha256:8a5d146eed3431f386cc232e7aeb607520c8aa09f979dbdae8beda8f546ba7de

Observation a2a6e250-6339-4122-b15f-b4bac850f715 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.615902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.615902Z digest=sha256:d44c1b5de1f08b7a4c8d7e072f8b783d6a38996785bff89191f983ab901b22a3

Observation ab58eaea-b3aa-4790-bff5-1bc003501ea7 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Opt: Open pre-trained transformer language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.707629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.707629Z digest=sha256:ea4d4336e62f5b2bc9faaabd2a04b5c5d58e7ab49cb7d47fb97eaf45e939455d

Observation 543d0be1-ba38-436c-9a08-fd300a83dc8a · outbound

This paper cites Qwen2 Technical Report.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.849854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.849854Z digest=sha256:1e2e561fa378fa65964a9ab8e9b5473f608eadf394d70fbd5cde998e0695bfb9

Observation c589949a-a181-4c18-b42c-a44f50ab1084 · outbound

This paper cites Ai and memory wall,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Ai and memory wall,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.961186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.961186Z digest=sha256:aef4b0d570e598b11c7926309d98084bc7d53659d056a8db8205db9beb2d19bd

Observation faa8d4de-c4d2-4dc7-919c-4814191e17ec · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving OPT: Open Pre-trained Transformer Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.775813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.775813Z digest=sha256:1989896ff720e3af5da40c94c2b22b622881c566a3fab999809d871fcde9c855

Pith citing papers

No inbound Pith citation observations are available.