Pith. sign in

Paper Citation Record · LEDGER

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling

As of 10 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2608.01891.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01891 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:49:27.634595Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c8b61952-edf6-419b-8c18-0766b3e15927 · outbound

This paper cites Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Taming Throughput-Latency tradeoff in LLM inference with Sarathi-Serve,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.396967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.396967Z digest=sha256:0854d0bcec9c0eafebea701caf3afa64282a4f443a65ecec61eca499c4018edf

Observation 6f6ae754-fee4-4a2b-a942-2cc2210b866c · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.474198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.474198Z digest=sha256:6bd6d0b96409e5eee5ba285db4c29781497ecc73b0a27b110a709eb0472de23a

Observation 5146295f-fdc9-4792-9e52-37b510fcd8fe · outbound

This paper cites Azure Public Dataset,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Azure Public Dataset,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.548021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.548021Z digest=sha256:859d878f0bf453a3338178989cdcbd088c832e8ad751a68d85e8dad7052a930b

Observation a0f222f9-27ed-4bad-b2e1-8f9b7fd4991a · outbound

This paper cites Biscale: Energy-efficient disaggregated llm serving via phase-aware placement and dvfs,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Biscale: Energy-efficient disaggregated llm serving via phase-aware placement and dvfs,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.681782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.681782Z digest=sha256:2bc07700e2b6c645d1cc69b8a73ee698049d907d87818d2c8a7ea259d772af72

Observation dd8891bc-fa31-4f8a-9b9d-e67073499dcd · outbound

This paper cites Reducing the carbon impact of generative ai inference (today and in 2035),.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Reducing the carbon impact of generative ai inference (today and in 2035),

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.836600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.836600Z digest=sha256:d5abee7abc0177e9f7b673700623e5b07c6e6f3eae258d4057dc33bf1cd0f1ce

Observation 21e9e55b-1e0a-40e3-8c73-ad1defc5eff3 · outbound

This paper cites Copilot,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Copilot,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:21.970638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:21.970638Z digest=sha256:96fedca3aa8f2fff33a42716c8f97c0ea930a66d5d70d7586eabd9cc41afa59a

Observation 7adb02b2-d8de-4fd5-82cc-15ff0deabe3f · outbound

This paper cites Flashattention-2: Faster attention with better parallelism and work partitioning,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention-2: Faster attention with better parallelism and work partitioning,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.121170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.121170Z digest=sha256:a16970a3e0aab3781f4bd7975f8cb3d71329da1a5c1d0c66299c6db4749fdbef

Observation 1a4e0ba6-d21e-4b69-8d7d-915c9b3fcfe3 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention: Fast and memory-efficient exact attention with io-awareness,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.256771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.256771Z digest=sha256:f67d8f9b136d079dcc5b8ae23812b90116bc9ab28a43a06e27883c01a70561b3

Observation b169a158-7018-48c4-90a4-f1e0e3fbd971 · outbound

This paper cites Context Caching,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Context Caching,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.399419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.399419Z digest=sha256:1ca15b5461d7e03c4190ada0b5740617ac7609445ec217f56f5445ecceb2a8e4

Observation f9d0b40c-a746-425d-a2eb-54a225b0ff58 · outbound

This paper cites FlashDecoding++: Faster Large Language Model Inference on GPUs.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling FlashDecoding++: Faster Large Language Model Inference on GPUs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.543190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.543190Z digest=sha256:4eb0f4895003120bc34758a7ae511e021a6d9b4c5a89ec13a834705035220244

Observation c08e005f-87c2-4338-99d1-66eec8b4dc53 · outbound

This paper cites Smoothoperator: Reducing power fragmentation and improving power utilization in large-scale dat- acenters,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Smoothoperator: Reducing power fragmentation and improving power utilization in large-scale dat- acenters,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.686898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.686898Z digest=sha256:a206a661c64afd0aebc484d1970baff648e3c3521ac70519a787b33377db9799

Observation 2b17282e-2cbb-4913-896d-4eaa09d08b5c · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.840559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.840559Z digest=sha256:1198c68155b7ed1e09423b669740a02e16d3fd8bf19c566625b2b45c8fad23f6

Observation 7d2fc877-c3e9-4a5b-9681-8eb28e071b5f · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:22.951059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:22.951059Z digest=sha256:bbd105747f79fe80dab13e1c2ab899a4e09d752ebe15b647a534128b5e399a5a

Observation 457c84e7-5678-48f6-a755-cebf572dc9d2 · outbound

This paper cites throttll’em: Predictive gpu throttling for energy efficient llm inference serving,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling throttll’em: Predictive gpu throttling for energy efficient llm inference serving,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.116685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.116685Z digest=sha256:ec1ad2e2c9f6f9b0de450a0a1069715c298ff5fa31f86e6f2e9ee07f924a01ce

Observation 53e7ff0c-8bea-4d4e-9641-10d6b95e32f0 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Efficient memory management for large language model serving with pagedattention,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.291029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.291029Z digest=sha256:3fdcf7da2a2e3e81c7fd4d8ef8b8ffbbe54909806e388618a14f008fcf450ca9

Observation f7303883-8393-45f5-a34f-296637f6ce26 · outbound

This paper cites Greenllm: Slo-aware dynamic frequency scaling for energy-efficient llm serving,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Greenllm: Slo-aware dynamic frequency scaling for energy-efficient llm serving,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.545949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.545949Z digest=sha256:706ad8b18d9bdd81912f08d62e92e268e9e4469197b5197d81ee810ab6f94c4c

Observation 794e38b3-7266-455a-84f7-dca308ad216b · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Lmcache: An efficient kv cache layer for enterprise-scale llm inference,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.710023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.710023Z digest=sha256:e23422c79df0060b2ed32c56411584b239dcf6cdd9551a35d6a51a7f7de6205a

Observation 9ef64e5d-7e77-42b0-8889-997a0ab070de · outbound

This paper cites The Mixtral-8x7B Large Language Model,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling The Mixtral-8x7B Large Language Model,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.826123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.826123Z digest=sha256:8e4fb14966a55fb6e450fccf205cb7e38440f0b66e5799ae4f5590a10fb78630

Observation 7a549a92-6e5b-4dda-bd38-c611ca53b34b · outbound

This paper cites NVIDIA NVML API,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling NVIDIA NVML API,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:23.994646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:23.994646Z digest=sha256:e8c04280ef5104a4f18f219e7be9abf65848c62c0e2d2bf4637f38d8af5851ff

Observation d43d3f0f-5404-4a5d-a9a7-df170b12c193 · outbound

This paper cites TensorRT-LLM,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling TensorRT-LLM,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.167873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.167873Z digest=sha256:55f27c079df5732ed2c00f23eb14359efd2cb6231ca7535b44054e09b1f10485

Observation b59d7b1f-d478-4908-a202-92558e179c5f · outbound

This paper cites Chatgpt,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Chatgpt,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.333635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.333635Z digest=sha256:333e13202244016750266fc41003c53ad692f501706c171e8bde4fb65115ebe9

Observation 58744b07-d77d-4831-a327-7c569edba536 · outbound

This paper cites Splitwise: Efficient generative LLM inference using phase splitting.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Splitwise: Efficient generative LLM inference using phase splitting

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.556527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.556527Z digest=sha256:e9ca9b82f9f84206f4d68deaa1ad2939db6bfa7cc4b9beac04615ca3b72af78f

Observation 5ab4d207-f061-4a53-b638-09625136db28 · outbound

This paper cites Characterizing power management opportunities for llms in the cloud,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Characterizing power management opportunities for llms in the cloud,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.764905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.764905Z digest=sha256:d4d38bc901b2f4c9e6fbbf4d5bba761b134018ab4d9044e227f1b16c093676b4

Observation 54002eba-c663-49c7-bb85-6036f9d4ec6c · outbound

This paper cites Scikit-learn: Machine learning in Python,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Scikit-learn: Machine learning in Python,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.900705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.900705Z digest=sha256:3c6996dff8cad9ff7123b485cfd900749e0795a564daf70a4cdf701de02852a8

Observation 47f16d29-84ed-4854-a1d6-e6d4afe4c52b · outbound

This paper cites A study of generative large language model for medical research and healthcare,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling A study of generative large language model for medical research and healthcare,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:24.994295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:24.994295Z digest=sha256:962f5e6b3ab3ed40746d496c7212231dccc2ab56814b8fa3ee07d99d4eeb2beb

Observation 5ab04deb-8913-4d50-9b27-0d582d93e81b · outbound

This paper cites Mooncake: A kvcache-centric disaggregated architecture for llm serving,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Mooncake: A kvcache-centric disaggregated architecture for llm serving,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:25.240563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:25.240563Z digest=sha256:da8b40202120a5faec5ca4cd7dbd0f40b4df815d592a84477b7664777c11b19d

Observation e9ebd65c-c8b8-42f3-8606-daea37da5e80 · outbound

This paper cites Accelerating retrieval-augmented generation,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Accelerating retrieval-augmented generation,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:25.412580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:25.412580Z digest=sha256:8962532cd355cb4cc7acf8d8cbaca65d0b02c4ea01ac8a6802f5f8de9f0160ed

Observation 3cded0ae-83a2-4c83-b703-5000c6474660 · outbound

This paper cites Flexgen: High-throughput generative inference of large language models with a single gpu,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flexgen: High-throughput generative inference of large language models with a single gpu,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:25.582523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:25.582523Z digest=sha256:d9aace3db1a76d029341c371e26dbcb03c150e05f56bbb7a5132d392556a0867

Observation 7c39d644-88c3-475f-8410-d9ef367d75cc · outbound

This paper cites Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Towards Greener LLMs: Bringing Energy-Efficiency to the Forefront of LLM Inference

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:25.678359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:25.678359Z digest=sha256:4619bf3fecf9cca73a94943f399eb6ee6dd94a93b3440622e50a1d32a3d7c821

Observation 9d2e6577-78f2-483f-8cd8-6c321991c303 · outbound

This paper cites Dy- namollm: Designing llm inference clusters for performance and energy efficiency,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Dy- namollm: Designing llm inference clusters for performance and energy efficiency,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:25.845086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:25.845086Z digest=sha256:e592c124f327db55e2e6ee906e3c82d121b74a0f63fc8b14666aac9348960a2c

Observation 84f707d1-3c7e-4b6c-9620-10046c86d5c1 · outbound

This paper cites Qwen3 Technical Report.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Qwen3 Technical Report

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.105281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.105281Z digest=sha256:c6794ff310726655b106806f40b6ee2055999be36fb377bbec0a31a3c86c4f78

Observation e741e68c-649e-4678-8028-0a09900ed3ec · outbound

This paper cites an unresolved cited work.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.278741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.278741Z digest=sha256:c61ae3fa8c3a838c29170df425f3b0cde578ad99691414e14a625a7b897b978c

Observation 32477326-cf53-49a0-bba3-40945c07fecc · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Fast Distributed Inference Serving for Large Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.444410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.444410Z digest=sha256:64e9b44c5c3c2dc37a17e602bf5e4ea9d397fc1962ad45aadfc9a4c0fe15217a

Observation 68d6a954-3afd-4f8a-9844-63303edc836d · outbound

This paper cites xdeepserve: Model-as-a-service on huawei cloudmatrix384,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling xdeepserve: Model-as-a-service on huawei cloudmatrix384,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.618236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.618236Z digest=sha256:cbccffbce6ce4290581fc05eaff75374379cae59cf867ca4945c91a9b77134dd

Observation 4cb5a811-269a-4111-b665-e4ca18f0ba36 · outbound

This paper cites Orca: A distributed serving system for{Transformer-Based}generative models,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Orca: A distributed serving system for{Transformer-Based}generative models,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.807560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.807560Z digest=sha256:17326c0068b6f5e52408d1aed5595962200f6377e67f17f35c0fc10ab7f42387

Observation 613bd53a-614b-4575-9df9-2b3852c3a5e1 · outbound

This paper cites Flashattention-4: Algorithm and kernel pipelining co-design for asym- metric hardware scaling,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Flashattention-4: Algorithm and kernel pipelining co-design for asym- metric hardware scaling,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:26.972179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:26.972179Z digest=sha256:d44438a6948431dc397ba43fe191db936ff15b7d639965b7d63997e4d72c6712

Observation 130dfaaf-46ba-4f50-abe4-03c045b3c6c3 · outbound

This paper cites ZeroMQ API,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling ZeroMQ API,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:27.145577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:27.145577Z digest=sha256:5bf3f2dc3fb3059458df7e753091592d952df208b8aedb36d0caa8fd7a1183fa

Observation 54ee6a5e-891b-4c64-bad8-269bd69d9b41 · outbound

This paper cites SGLang: Efficient Execution of Structured Language Model Programs.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling SGLang: Efficient Execution of Structured Language Model Programs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:27.314415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:27.314415Z digest=sha256:41d574bd6f2aa3af14fd87c3beff8dc54f9448ae014b83aeda1e9358b02e4f47

Observation 3c37b2b4-f27d-4ada-8ab9-0926f0f264ab · outbound

This paper cites Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Distserve: Disaggregating prefill and decoding for goodput-optimized large language model serving,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:27.475602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:27.475602Z digest=sha256:ae7861198d1ebf32a7fed9e0c91e9080e72dda561f22a53630e8a08d6bc5a4b0

Observation ab2e0bd0-1e48-40aa-af6f-10cf8bedd45b · outbound

This paper cites Megascale-infer: Efficient mixture- of-experts model serving with disaggregated expert parallelism,.

Energy-Efficient LLM Serving via Disaggregated Attention--FFN and Flexible Frequency Scaling Megascale-infer: Efficient mixture- of-experts model serving with disaggregated expert parallelism,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T18:49:27.634595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:49:27.634595Z digest=sha256:1833bba0514d1fd674be03a9931f35651520c1acff073156edb09ce372a98613

Pith citing papers

No inbound Pith citation observations are available.