Pith. sign in

Paper Citation Record · LEDGER

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving

As of 7 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2607.29575.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29575 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:22:54.961186Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1283f76c-0bbf-4207-80e6-721c88cf4e48 · outbound

This paper cites The rapid adoption of generative ai,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The rapid adoption of generative ai,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.621032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.621032Z digest=sha256:d85e7de237533e089a046e95fd0ec06270fed07ce47fc574e594a5d0492de23e

Observation 71750f8e-e65d-4dfc-8fc7-5c8e831faba4 · outbound

This paper cites The adoption of chatgpt,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving The adoption of chatgpt,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.687853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.687853Z digest=sha256:404643356aba82085fcd5a5fb7c745ac2b6a17d655498604fffa28a34e53805b

Observation 5e3aec83-1a24-4388-adde-a9bb58b1fdf0 · outbound

This paper cites Quantifying large language model usage in scientific papers,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Quantifying large language model usage in scientific papers,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.737457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.737457Z digest=sha256:b88d55babb1941d5d48ca3d9419b570afacdaa1e942dadc27566311248c2f5de

Observation d7a31b72-eb9d-4ed6-8138-bb0287ef2850 · outbound

This paper cites Llumnix: Dynamic scheduling for large language model serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llumnix: Dynamic scheduling for large language model serving,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.844613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.844613Z digest=sha256:b246f0ab7b28b9160aeb7ac3bdd81ef251f0b2f9ab0abcd8e6aa247b20f753e0

Observation a627426c-a4f4-4206-a394-04d805de12ae · outbound

This paper cites Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Sageserve: Optimizing llm serving on cloud data centers with forecast aware auto-scaling,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:51.935319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:51.935319Z digest=sha256:acf75b3adab5e56d878344860baa00e78fbcf18979919eccdbca58adda6b169a

Observation dc0fee93-9e4d-46c1-8b28-cfa5f828b8ae · outbound

This paper cites Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Large Language Model Inference Acceleration: A Comprehensive Hardware Perspective

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.032744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.032744Z digest=sha256:235896f4715c3103d1a5cdb235fb1a15507f76c167a0fc5cd211036b6c3bf7a7

Observation 3d7e53ee-93d8-47cb-a734-76e6e4e5c880 · outbound

This paper cites Efficiently scaling transformer inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficiently scaling transformer inference,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.082197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.082197Z digest=sha256:a42b94f596a5bf87296f52d98e6c0ba60d6243987faf61085cace90200e45a2e

Observation dbd5fb7e-2239-4859-89da-77361ba9beb1 · outbound

This paper cites Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient llm inference: Bandwidth, compute, synchronization, and capacity are all you need,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.167267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.167267Z digest=sha256:53a80379d4f6a4ed732c73c62bef2cfedcc89d1837074091d10eef930f975800

Observation 8401ac48-13a1-42e6-96f7-d75a09241a32 · outbound

This paper cites Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mind the memory gap: Unveiling gpu bot- tlenecks in large-batch llm inference,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.228945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.228945Z digest=sha256:6e0b00200490fbf64a894466ba14cf615097c61a1fd48522fec69afcc8cc32a6

Observation e9be6525-2ed8-4deb-baa1-20852599966e · outbound

This paper cites Llmvisor: A real-time latency attribution model for multi-tenant llm serving,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmvisor: A real-time latency attribution model for multi-tenant llm serving,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.312448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.312448Z digest=sha256:44fba6a1e0d359d25eac062fb4010d9c8a02d92220c4d8a33a9d9577ec5b5044

Observation 169dd382-9cd3-48bf-8c43-d32d50983289 · outbound

This paper cites Predicting llm inference latency: A roofline-driven ml method,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Predicting llm inference latency: A roofline-driven ml method,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.386461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.386461Z digest=sha256:bb0293c779e93de93a64bd24513bc94a80f24766fbff2d15b47cef79f0fbf065

Observation faa6fa12-8937-465d-87bf-2c10a68eb48a · outbound

This paper cites Language mod- els are few-shot learners,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Language mod- els are few-shot learners,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.529800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.529800Z digest=sha256:9d12e399d125b071fbe62da86d33104b9d7700a181c282540097944493559472

Observation 06875560-041e-47bd-bca5-4e0a529eaf44 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLaMA: Open and Efficient Foundation Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.616647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.616647Z digest=sha256:e84d06d9290d0c90f79c7c31f541a140a7e0f9069fba84f5783c3f28289f1755

Observation ab551368-fb6b-46ed-b8d7-789c77352651 · outbound

This paper cites Attention is all you need,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Attention is all you need,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.703801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.703801Z digest=sha256:e6d34884686bf2d4bcbf0f04325d6ac74fc43de6f7ba983ba750744563d151bd

Observation de2897af-d4a1-4f69-8e86-b899daadc481 · outbound

This paper cites Orca: A distributed serving system for transformer-based generative models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Orca: A distributed serving system for transformer-based generative models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.769107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.769107Z digest=sha256:3481c02417f75acc83ac630c2e97d4821e02208e2bd0388779c67d6d4cbd4e71

Observation 56f5eeae-c9e5-433a-b6c4-22b4f42723c9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Efficient memory management for large language model serving with pagedattention,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.843887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.843887Z digest=sha256:15add55d9846d82c6b3a8bc44e6a2cf9389e9fa5234b7a542977ba57d3d3b086

Observation 3aa6c78b-0554-4beb-a73e-df1f537fc2e2 · outbound

This paper cites TensorRT-LLM,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving TensorRT-LLM,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:52.928323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:52.928323Z digest=sha256:58e78fbfd6ab25860b5a7b6a5833a63d6443e87b33e5bbc704b86b1b993301fc

Observation 94854dc3-b780-4314-b2a6-ef0ecc871228 · outbound

This paper cites DeepSpeed-MII,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving DeepSpeed-MII,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.030125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.030125Z digest=sha256:08aaf69c787980011bf19dce4cfdd4a27b0aefce1346436770604b75d1eeeb73

Observation f142fcee-c690-41b7-a722-d974ca81131e · outbound

This paper cites Slora: Scalable serving of thousands of lora adapters,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Slora: Scalable serving of thousands of lora adapters,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.117570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.117570Z digest=sha256:5a997144c698a119a028b18f2b9d89fe908f767bc49624bdbc23b61f4404077f

Observation a133c8c9-ab5f-42b6-a05b-570a8d2e5d44 · outbound

This paper cites LLM Inference Unveiled: Survey and Roofline Model Insights.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving LLM Inference Unveiled: Survey and Roofline Model Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.202838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.202838Z digest=sha256:e4da40a06a516dc84f6d026e13674086dcc38c4fc2bd8969e899baf8bb68198a

Observation 46ded418-0c69-4282-b2fd-5ca8d11f8edc · outbound

This paper cites Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Taming{Throughput-Latency}tradeoff in{LLM}inference with{Sarathi-Serve},

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.305385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.305385Z digest=sha256:109c7e358baf0344cdb6e39f88eb4f31e8520a8d8e6cf87fee152b611b543d6b

Observation e556d816-ff08-40d6-835e-f20286af23c8 · outbound

This paper cites Fast inference from transform- ers via speculative decoding,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast inference from transform- ers via speculative decoding,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.428587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.428587Z digest=sha256:2489987af9138f609f0cb5239897aed19ba24a64dc33620a2b77befa3f87a56a

Observation 11947e0b-a77a-40a4-9b7f-51eeb4fa4fc2 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Accelerating Large Language Model Decoding with Speculative Sampling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.550502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.550502Z digest=sha256:f8f3cc482ffbd1de806d4ef94e5cb674145c80e4aa88359070956dba0d6f35a0

Observation f0c8d29b-ce4a-4423-b2d6-fc9afdaae304 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fast Transformer Decoding: One Write-Head is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.668526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.668526Z digest=sha256:a65b6f857d8430f9501c04cfaff0660b9e98c1ca28482049068a1abb35a880e3

Observation 4fff4c06-15c0-4d6c-82a7-e8a0c5e61c97 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.767380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.767380Z digest=sha256:3f3df08c303b1cef6a6365b6a495cfbae708160903a618b9c7a99dd609fa36c8

Observation 8a9d34a5-5e25-41b3-b4f1-f9dde91eb23f · outbound

This paper cites Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Awq: Activation-aware weight quanti- zation for on-device llm compression and acceleration,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.847735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.847735Z digest=sha256:f110c578efae5ff875df82ce59cd2d2f85406419bfd206d2bed3ea285a6e74a0

Observation 6f35459a-2052-4824-bc59-7fc008a098fc · outbound

This paper cites GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:53.915456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:53.915456Z digest=sha256:13feaea1bea81ae0af7e39aaa7d688c8eef9e76c884f3995c335cfe2d168ce17

Observation ec390e4f-1a26-45ca-b8ef-c96e3928b659 · outbound

This paper cites Vidur: A large-scale simulation framework for llm inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Vidur: A large-scale simulation framework for llm inference,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.024548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.024548Z digest=sha256:4b4ff08b89a8bcf70570758e549ea64b392ffb38aa7d4d799287f8ac6fa212de

Observation 9474e2ac-8847-448d-a2c5-33784756ba3a · outbound

This paper cites Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Demystifying AI Platform Design for Distributed Inference of Next-Generation LLM models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.136193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.136193Z digest=sha256:8eb118c07a9d51903c93a491caa0c00ab48dc059bc2331e9486e0d82640f93ad

Observation e3a09dfd-6bd0-4c5c-b9c0-dad1f0f4702a · outbound

This paper cites Llmcompass: Enabling efficient hardware design for large language model inference,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Llmcompass: Enabling efficient hardware design for large language model inference,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.209163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.209163Z digest=sha256:d4e101bd2141f80f4311428de6e7e93c6bc2acc99e283b953a1a28f2173d49bd

Observation 8d956b91-494c-4617-83d1-4b357831ad48 · outbound

This paper cites Amali: An analytical model for accurately modeling llm inference on modern gpus,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Amali: An analytical model for accurately modeling llm inference on modern gpus,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.321060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.321060Z digest=sha256:8a92944d2d88219390cd2b0418d3789465f191df4ba6c2d25cbda3e51ecbf760

Observation 46db8890-3e1f-4963-9f18-6cc14c0f8fef · outbound

This paper cites Fairness in serving large language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Fairness in serving large language models,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.393342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.393342Z digest=sha256:0963e3b25b3041ac2cfb0877a518e6edf5012005faeb59b4a7348d74d36e6bf2

Observation ef1c6471-f486-4f1d-a3b2-dc42413c0c53 · outbound

This paper cites Clean sharegpt dataset,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Clean sharegpt dataset,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.458412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.458412Z digest=sha256:3101426700a8da75e96e9e6fb0173b9ea6a646a9333feb7d956832976f647518

Observation 6b314849-88b5-42a6-910a-381b7b4b23f0 · outbound

This paper cites Mistral 7B.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Mistral 7B

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.533081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.533081Z digest=sha256:81af38a33d855a4ac5b3d24903aab8b3c235976a1abb9294da35a1c7cb85c5bf

Observation a2a6e250-6339-4122-b15f-b4bac850f715 · outbound

This paper cites Granite Code Models: A Family of Open Foundation Models for Code Intelligence.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Granite Code Models: A Family of Open Foundation Models for Code Intelligence

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.615902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.615902Z digest=sha256:8b93d9ca3dc463c6d5a12aae4a7de7cd65e29480c6a6474bc6c6f1194cb7b820

Observation ab58eaea-b3aa-4790-bff5-1bc003501ea7 · outbound

This paper cites Opt: Open pre-trained transformer language models,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Opt: Open pre-trained transformer language models,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.707629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.707629Z digest=sha256:5852cf9f54af82c4d7037b5c38b36bf4c28ae29ddf846294ad84843a16da8dd5

Observation 543d0be1-ba38-436c-9a08-fd300a83dc8a · outbound

This paper cites Qwen2 Technical Report.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Qwen2 Technical Report

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.849854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.849854Z digest=sha256:1126a57b4288a134fd43dc267736c78831d23093d9eae7d4718287cd8b0c1b20

Observation c589949a-a181-4c18-b42c-a44f50ab1084 · outbound

This paper cites Ai and memory wall,.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving Ai and memory wall,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.961186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.961186Z digest=sha256:9b50668c97077562b4dfc11e06f7bb83a73ac775b386b826bddf298e00386b1c

Observation faa8d4de-c4d2-4dc7-919c-4814191e17ec · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

SLIM: Saturation-Aware Lightweight Performance Modeling for LLM Serving OPT: Open Pre-trained Transformer Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T04:22:54.775813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:22:54.775813Z digest=sha256:48304ddaf7f0fc92dd8fa9b923a985d896fc52809d51b28c7e812a435c653cef

Pith citing papers

No inbound Pith citation observations are available.