Pith. sign in

Paper Citation Record · LEDGER

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding

As of 11 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2606.29207.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.29207 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T02:54:11.777615Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact14
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe7376e2-ea5a-4d37-add5-09056dfb6d34 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:400912f290cf462597655cd95bae91df44e809caaee38bbaa66b050ed10fdcbc

Observation d25d1953-9c09-4cf4-b309-3134ed018d6f · outbound

This paper cites CONCUR: High-throughput agentic batch inference of LLM via congestion-based concurrency control.arXiv preprint arXiv:2601.22705.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding CONCUR: High-throughput agentic batch inference of LLM via congestion-based concurrency control.arXiv preprint arXiv:2601.22705

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.254956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:3956c87815cea74478780b4ed54fef2eb523af787ad3d9af18d43b8fb6f8efe1

Observation 41dcbe9c-0c08-4e4a-8bfc-9dd7f5679af0 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:a6e6897aed122bb54449a4c3a6ad8a07ce294d0867572991e00d74b9906fc507

Observation d67e1e4d-a683-4286-9711-5432ebc161c2 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:59c794d8d3e5e1fbd03d0390cf13b0735328556b7109f2fe92560e2e37eeae09

Observation c54d6505-605c-4cab-b07d-925fb5e10c9e · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:a59413f43216262b896f07ec27a5bb1927e3128e973f30ae7143b059f210866d

Observation fa50a89c-e5bf-42d4-93b2-81e8cde0ac98 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:2aff2800e76c83b68469bfb5f1347f27cb84ca08bf785e624274ab716a387fc3

Observation f9c73c01-3186-4420-936d-086af705281e · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:4009367830fc74bb546ce1fa727370c03912af6b835ad53394279973b1da019c

Observation a5a512c8-266f-4b28-bc4b-4202f6ce6f64 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.212475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c79b4968502c51def2c240789609d223a4cbe9fd85e6e4df1a97c72d8ffd9926

Observation 997d0580-e239-4e76-b7b6-b1f78961ae40 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:ff33711461b4810b4f721293216c2285651436fde860b43986b71c07346b3bb1

Observation ef9aefc8-c8cb-4722-b9f5-3ab05fb142b0 · outbound

This paper cites MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding MemServe: Context Caching for Disaggregated LLM Serving with Elastic Memory Pool

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.223576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:407b7f09c7ea3d6b91a020db9f43cd10142edb434aa7252df48c771d64503d33

Observation c6592f79-3a9c-43b3-88c1-430de1a4073c · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:acffab679af48236c865f822907688e43697a9229f15beda0ebd4ea4291ff5df

Observation 54628fc8-b21f-4a37-a66b-2aea8d0c9c70 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.237137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:d6e365b14822cc050534cccd9979e07d93eb61d3349ffb01abc934c221fc53ac

Observation 66e6b0dd-2217-4ded-8784-54051512ebd8 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Gonzalez, Hao Zhang, and Ion Stoica

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:7248dba167b9fde59e4bacef94fd044baf7b9c652b048e188da758f65acd6628

Observation 35c3e917-cc1a-4083-9fb4-bdabff066576 · outbound

This paper cites In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23).

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP ’23)

Reference 14

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T03:04:14.265268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:6c8575c513b48dc039ce1107166cb69f2fca64e9c4b9c90362b2efc5a9d4582d

Observation f1d38012-a07f-4c57-9e7d-ee67d6f3bd26 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:31052bee18525e5201f9294f4175b2b96199b1959ce3c3cfe94066095828bc87

Observation 0c43c7c5-d02a-4706-991e-1671fce3e692 · outbound

This paper cites Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Infinite-LLM: Efficient LLM Service for Long Context with DistAttention and Distributed KVCache

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.208841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:30300ea540d45d1f0c8ccc2e46a1690a715f04a9917e2af07c6275e094653197

Observation c1967422-3b95-4e7e-a473-9e005a47b312 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:cf5c7a639a41da3eea22e4ec40fa7a57801bbb7402d1ba50bc338e921cd8ecb4

Observation 4e0773b8-588a-4436-879a-bfb7559130d1 · outbound

This paper cites DeepSeek-V3 Technical Report.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding DeepSeek-V3 Technical Report

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T03:04:14.260714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:0a6f5eb5821a04b0213a04c4dcab208c5898bf0f6203557dc779a3bf5aaac809

Observation a13546a9-8d61-4240-8055-2ed2fa3a89a5 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.219523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:efcc8c5917114fdfe502e91b39aae52be99432441df5376459cfd3841d7f69c5

Observation 3d97db2e-18fc-40d2-82f1-6101e92a8ea3 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:04:14.180328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:0a624b4950a96a3e3ec35a0eed1190591cb133a954a29047c1bb891206c8126d

Observation 6ad5adf1-b131-4df4-b2a7-108f7e5543d1 · outbound

This paper cites Online normalizer calculation for softmax.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Online normalizer calculation for softmax

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.188442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:4064d5f4c2ba77baea9592522e953dc556b3e8b300ee8aaba14dc93b0d0f9885

Observation d3bea937-a7f7-4853-878c-7318feead93e · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:4900b6637a920d04f619f3427fd50ed4b6bb05f58cf34416afee898a852c059f

Observation 9be2a688-cb7b-4fec-a2f5-39d79eab3378 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c223eb95eb0febbfcbc43fbf29f21853a7fa13129aba490e881c412cc9187dfc

Observation 949f75b8-5c44-477d-b819-b7363f75e0a9 · outbound

This paper cites MemGPT: Towards LLMs as Operating Systems.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding MemGPT: Towards LLMs as Operating Systems

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.176459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:8a065187c4c4c19c82e88936c1224311e71e468b5745e11e9a95710933dcb756

Observation 12256fb6-6ec4-4136-8c02-12093833f48a · outbound

This paper cites Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Lev- skaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Reiner Pope, Sholto Douglas, Aakanksha Chowdhery, Jacob Devlin, James Bradbury, Anselm Lev- skaya, Jonathan Heek, Kefan Xiao, Shivani Agrawal, and Jeff Dean

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:04:13.705243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:00e18bffa30c819b7d72801606c08732d02da26bd1f8dfeb898c75a9f9c4792f

Observation dce82fe2-6899-4a76-ab18-64fb9f9c6ed0 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c815197d11dfba02cf49554bdcff4a38e923ba5b3e1aae73b03c197c5c1857f1

Observation 7987c3ea-151d-40bc-a969-6c3bbda62804 · outbound

This paper cites https://vllm.ai/blog/2026-05-06-mooncake-store.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding https://vllm.ai/blog/2026-05-06-mooncake-store

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:3128185d37215a232735cf0fe24e3e3392e77c37f5f6ba701cac87ec1b367302

Observation 4691849d-b3a1-4113-8de7-04fdc8835d0d · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:6949d434956a983540ae750332ecde9d388f7f13b8c54ef165e18fa951dc3177

Observation 8803a508-7e19-4ba7-811a-5f1fe3f14b7a · outbound

This paper cites Graham Lopez, Matthew B.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Graham Lopez, Matthew B

Reference 29

Resolution
malformed identifier
doi_truncated, observed 2026-06-30T03:04:13.700088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:6c2ecb0f1f0c57fa536577bd3b7e0d4a39229ef1736db858e8b3e7e28381a218

Observation ac1f5f37-f81c-439b-a554-f53a4a86c4e2 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:291e0945f9da38b6e2d8db3e8ed73bcc84b02d2ad3f41125eda62ebcbec17569

Observation be7a581e-0f5a-4451-b6c0-6a1410b0eb13 · outbound

This paper cites InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML).

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding InProceedings of the 40th Interna- tional Conference on Machine Learning (ICML)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:3d9767593a9eed238c8f419031a2cfd96bb713c1b5438dc011faa68c12fee509

Observation 89db038f-bc64-4bd1-b001-471606a21070 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.170976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c652cfce1531d88769df795c1e9cd2422a5e4bd77026d4370f27e6a924d6405b

Observation 3e56425c-2ba6-40b1-a5f4-ae240261e427 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:8c73ea3d514571c481d9ebaad5150fa57f7ef587e69eaac9820d6737b2907a2c

Observation 683ca576-4ec4-45a9-9f5f-cab250b9ccd5 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:e29324e7386e73849d78dfaf03ea35a0b0b77014aea13ffd281f7b6b9aa47bf4

Observation 58ecf3bc-65b8-4c89-a3fc-f269709113c8 · outbound

This paper cites In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP '24), 2024.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP '24), 2024

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T03:04:13.710550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:247dae57d741d713f9ad35c080d84e5a7100927d14cbc4d965246947f5497bef

Observation 8bd74438-9350-4275-8799-d942ac98399f · outbound

This paper cites Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Combating the Memory Walls: Optimization Pathways for Long-Context Agentic LLM Inference

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T03:04:14.196446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:b948586ab4dfa3557dddb8faefe06f2c81dd26f438987924e4b25ce872cff526

Observation 68e52c3a-1e87-489d-b804-53d02e36f54a · outbound

This paper cites Dualpath: Breaking the storage bandwidth bottleneck in agentic llm inference.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Dualpath: Breaking the storage bandwidth bottleneck in agentic llm inference

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.268305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:8c21fd61eeab844fff8e41198acb756ec7bf90646de9251ad1eb9e66dc8ed12e

Observation 72ba9c36-d2ef-4cd3-b4ad-99e46bb75ca7 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:13803553501ca5e92cb001eba25296ae23d50657e58fc81ae9743ffe074be1df

Observation 744dfe47-89ce-4be2-8d8f-d136776e890b · outbound

This paper cites Context Parallelism for Scalable Million-Token Inference.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Context Parallelism for Scalable Million-Token Inference

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.241549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:e3529001d4c37c90f659d7e76b41e5750bf58760df33dce9647cfe259cabd39a

Observation cd4cdba3-61b5-4074-8bc8-58af2c9433b1 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:06ba8c7a3100dc11686a8a885d5f8bc39885cca6b2cf5572c2759f296fd41d9b

Observation 81ff348b-798b-41ed-977c-18c2823e3782 · outbound

This paper cites LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding LMSYS-Chat-1M: A Large-Scale Real-World LLM Conversation Dataset

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.246323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c383abdecf3a990685ee7640f9b652862730d282aefb72149690aba3bd275de8

Observation ab2abf9a-5406-47af-a8cd-c6d589012baf · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:14d1087ad2d2532c6b4644cc571ebc3a4b424cbf7261779ff5718a3171621768

Observation c3686bac-6a97-44b2-aa8e-c88c48e7e763 · outbound

This paper cites InAdvances in Neural Information Processing Systems (NeurIPS).

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding InAdvances in Neural Information Processing Systems (NeurIPS)

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:0ed6d23f85acab5dbc23803ec691fef8b6766d70f54feddcebb5ff020d2dbcc0

Observation 7fb983a1-5fff-4a43-b385-770b646b6d69 · outbound

This paper cites an unresolved cited work.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-30T02:54:11.777615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:c7e3e3d05eb77f1f915bfa033fbe40ce0b5b59535535b009ca3658aaee5ca938

Observation b7a52402-1472-4cb6-9520-55da65adaea6 · outbound

This paper cites MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding MegaScale-Infer: Serving Mixture-of-Experts at Scale with Disaggregated Expert Parallelism

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.249359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:13a29a835229db98344101aabbf0ee3ae9d11cb9917a00544a330edddd123249

Pith citing papers

No inbound Pith citation observations are available.