Pith. sign in

Paper Citation Record · LEDGER

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

As of 9 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.06007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06007 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:08.569863Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7266694b-2167-460e-a9a2-c6a1461ddcae · outbound

This paper cites https://www.kimi.com/blog/kimi- k3.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.kimi.com/blog/kimi- k3

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.210099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.210099Z digest=sha256:e350f6bc27b6ad8827957eb5e0d020a477511ab2ae74467628ebabbe5e12ecd6

Observation c5640c60-ddb8-46c0-820b-289b6a800bd5 · outbound

This paper cites ServerlessLLM: Low-Latency serverless inference for large language models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure ServerlessLLM: Low-Latency serverless inference for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.215746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.215746Z digest=sha256:59ddb8e2f1ac82d6c98a5b42c9650798fa8b1774d2b5516d18ff9d5fb6b106ee

Observation d373c80b-505e-4110-9072-7ef34e468468 · outbound

This paper cites Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.220987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.220987Z digest=sha256:f8bec77a6c981a66cb6ffbbac3596eefe53d22a9d6f74353e94ec239530b0f5f

Observation 059a9267-99f7-4671-b484-7768bf48ee96 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.225722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.225722Z digest=sha256:701376f4ddf0d82c0ecf28719fe55d3fbef26f25058173cec0707718b2b8483a

Observation 7bb633c6-8d07-43e5-a398-f51152382d11 · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.230680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.230680Z digest=sha256:ea1046095f6f01d23cddd48f4a28952737bfe5d42ed5368842df7d5fec64f17f

Observation e4436a8a-eac5-4d58-a312-6f25afa6d915 · outbound

This paper cites BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.234958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.234958Z digest=sha256:9da5d85dd7967866976cc0a0d71e1eab6c6c56d5f799a8ffdf62e63c2e0dcc99

Observation 4c447d41-9049-42d6-a526-314478418a4a · outbound

This paper cites Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.239382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.239382Z digest=sha256:c619c601cf50e86c77f87d7e4541990380c72b7a30d913435dfae113d8d46540

Observation cff16d1b-78d6-4d53-afdf-1a465b2f75e0 · outbound

This paper cites https://lmsys.org/blog/2025-12-10-rfork/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-12-10-rfork/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.243347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.243347Z digest=sha256:5b891d2374b7814dd2ea1557ed13c99ef9773b8b9b657122080311e9b2911c95

Observation 9012b892-fc8d-46fe-9e9a-83c5ef0d592a · outbound

This paper cites Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.247723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.247723Z digest=sha256:28f3c41ef0bd452af4bcd18f33b6335da153b9b854f257c5b75b2145e90c8820

Observation 214672d2-ef54-4bbc-9f95-a8ecd3432886 · outbound

This paper cites Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.252135Z digest=sha256:45f9ad25f79bfe7bca17c5385cc9fb0b6422bf40b144ce38f0d6472d66ac1235

Observation 530077e3-3828-4a3b-8d88-67308194f45b · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.256927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.256927Z digest=sha256:51184e7f112c896c4712e4cc28374edaad3a88d4f30cbde00fdb826ee9ffb0e5

Observation 1bab6d33-92c7-4dee-8e00-57df7e7504a8 · outbound

This paper cites https://lmsys.org/blog/2025-09-10-sglang-hicache/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-09-10-sglang-hicache/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.261289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.261289Z digest=sha256:c86eb4b67abd8e992ad66213821cae26b721ef8167d5fcb8928687f5841aaa3d

Observation 42f290ec-4062-48e3-bfca-971efb32b231 · outbound

This paper cites Stateful large language model serving with pensieve.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Stateful large language model serving with pensieve

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.647674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.265567Z digest=sha256:d0d558a6282803341abd05da0f30b8716fd830f63ed87fd0fd36038b78114d02

Observation 9194a7a9-2f73-4e9c-b28f-1f56071a2766 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.635789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.270101Z digest=sha256:d8cffe0a770908db2b6b128034727b56a4e33071b057c33f58dab23ac6449643

Observation da5d784e-d1e4-4906-82b2-e611454f3ffb · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.624659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.274263Z digest=sha256:553795bf411876d0dcb146b274f0aab2f307724f39c6eada992e9b823c163435

Observation 7b20c028-7b9f-4073-9b77-231851ecc7da · outbound

This paper cites {ByteCheckpoint}: A unified checkpointing system for large foundation model development.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {ByteCheckpoint}: A unified checkpointing system for large foundation model development

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.613146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.278745Z digest=sha256:5e8cf7195b39ea478ec2c622bbdb00dd48562c70b0cd938326e6b7f64e617f9b

Observation 4662b785-a21c-4aa6-a755-bec13fc799eb · outbound

This paper cites https://github.com/MoonshotAI/chec kpoint-engine.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/MoonshotAI/chec kpoint-engine

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.601572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.282891Z digest=sha256:2d7e0fac9123ef6da50a160e229e16ff06e4bc0676945a9c592d4ab1037ea706

Observation 31c3bf0a-5caf-4897-8dfd-3bc84f726e13 · outbound

This paper cites https://vllm.ai/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://vllm.ai/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.589655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.286808Z digest=sha256:d96ae0f9f1a8c7e4b79951ba8def1af076d87dcfd23ee533a5882d86139d98b4

Observation 4559eff5-182a-4fde-b2db-6844d6b4eef1 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.292156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.292156Z digest=sha256:0b3ae92fb25475760fd5ba6ce612b060e8582936274cc624dad09b911b0ef2e2

Observation f1a612b6-281d-44a8-966a-de5cd8b8260d · outbound

This paper cites https://nvidia.github.io/TensorRT-LLM/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://nvidia.github.io/TensorRT-LLM/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.571245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.296384Z digest=sha256:0d7eb9287747ecde25eab61bbb644407426e5bf925526df20ba2cc213cd0e1b8

Observation 88da8772-4852-4076-aeed-3d06e9d38c5c · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:10.559915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.300372Z digest=sha256:d500920a50cbccbb50736a6823dee3e74b4daba406727820fe6be92f88b14e78

Observation db842daf-c122-4ea2-933e-7cb220a4e955 · outbound

This paper cites https://redis.io/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://redis.io/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.546071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.304346Z digest=sha256:ef54a99f5f9cff22ca62013726e99e48573e7423f2f7f21e12c5e37b71aac9b9

Observation e67a7a63-1040-4cf5-a8c1-e12cedf1315d · outbound

This paper cites Vineyard: Optimizing data sharing in data- intensive analytics.Proc.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Vineyard: Optimizing data sharing in data- intensive analytics.Proc

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.534517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.308377Z digest=sha256:37720f73dd587378c50370cba60e625d35c3e0e6bd3bbead454d03fa080951f4

Observation 45e262c5-30d2-4c6f-be79-3d1fc74ab155 · outbound

This paper cites https://github.com/ray-project/plasma.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/ray-project/plasma

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.522032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.312295Z digest=sha256:f8796301c9e603afd3c3c2ebbe8a8263177b97eba9796cbff259631b18434926

Observation 8436dcad-3029-4bdf-bb28-e3e864b50b22 · outbound

This paper cites Ray: A distributed framework for emerging ai applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ray: A distributed framework for emerging ai applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.509673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.315974Z digest=sha256:bd5b79b2ad5c59f6442554f580eeb8bd6704f64120d762ce43fd73a61a4996e7

Observation 76dc0904-c3ee-4e49-a3cb-3e5b4a919e09 · outbound

This paper cites Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.320121Z digest=sha256:5341b01b3e3baf38a41409cc2f5b11a53288442d6b5183ad8829516517324e9b

Observation 0d236e8d-f09c-4621-b24c-a6a9d6cc73ed · outbound

This paper cites Serverless computing: Design, implementation, and performance.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Serverless computing: Design, implementation, and performance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.324667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.324667Z digest=sha256:4b57697c12ec9faa68ca121dc888760d9d3f71cc3698875b6eb81c46bb369b66

Observation b4725148-3cf4-48f3-bcf2-2e285e89f710 · outbound

This paper cites Infless: a native serverless system for low-latency, high-throughput inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Infless: a native serverless system for low-latency, high-throughput inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.478626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.328764Z digest=sha256:19f1a7774fa7651c4a2579e44f9615db5b8d967499bb284a1e79396472de0511

Observation 61af6026-a968-4494-9880-29304696fc9d · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.332432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.332432Z digest=sha256:fd9078d712b786d3606abb69ae760356d37ceefec60d149e3f488df67fd9112c

Observation d04194b5-e4b6-438e-8cae-de4d804b89f0 · outbound

This paper cites Language models are few-shot learners.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Language models are few-shot learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.336296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.336296Z digest=sha256:d5db0cb1d28e1bfc001a35343b68aedefc59008c578fb090f599f69a0e93561a

Observation 8c2c0de8-44a0-435c-8316-63e889a005a8 · outbound

This paper cites Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-08-07T18:23:09.340739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.339917Z digest=sha256:43afbc6e98885aeb7cc7484143d5f753cd5cf7c1a45099f78045d27248153eef

Observation 770044c6-56af-42e6-a0b9-808c69d64444 · outbound

This paper cites In USENIX OSDI, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In USENIX OSDI, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.450965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.344903Z digest=sha256:f8e7078c5b94fbb8b3a5094b51a5ff04ddd23988b918b3302941c311dbc7664b

Observation c0f22ded-c796-44de-a7ef-351c921e3fd7 · outbound

This paper cites Splitwise: Efficient genera- tive llm inference using phase splitting.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Splitwise: Efficient genera- tive llm inference using phase splitting

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.438704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.348995Z digest=sha256:c8be8e58fd1a05ec9d4d143c22d573b1aac313e92b383239be627b41be58ebfc

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:6ea47849040d73cc8f94cf5475d730061ce1a0e1e2fd79f50b04923f31417f64

Observation 60fe65b5-c7d9-437f-aa3f-c6cb4b316c79 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.357426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.357426Z digest=sha256:b1742166c29401c453c4638dbef755fe3545cbc2dfbfeff8d578584d00a04d77

Observation e3d048ff-38d0-4f2e-9257-ef5544fc095f · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.425673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.361591Z digest=sha256:ecab5786373eb01c875174a0a1e13509002db64cea2bfb085bdf8e8baffb5e15

Observation bcc27414-be25-4ed7-90ec-d706b5fad965 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.414206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.365396Z digest=sha256:de7e13f1d437a43284b84c3150744588e8b2fef14b63a1b7af84de100363a1fd

Observation 1a8aac5a-b716-4e1a-b453-3e151a60595c · outbound

This paper cites Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.401403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.369036Z digest=sha256:3a6e5b77b57a9c3a544a0d26234a1b0bce16b311781b035174bf42fa798fba61

Observation 573cad29-4ab0-4ad7-ae10-34ce8d5c21d4 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.373466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.373466Z digest=sha256:d03ea26131ee577e7e289f1e19c5f9c2582112c1396e8e85ed971013f176f437

Observation 254654b9-7c4a-46cd-95fe-2db36adfceca · outbound

This paper cites Dist checkpointing package.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dist checkpointing package

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.386334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.378063Z digest=sha256:45f33e45da804d307aa24348780b70778b2a8a3299205f9a890e6db2ec8f0b89

Observation 69313a8f-f463-45a8-ac16-3a10cb5bdf78 · outbound

This paper cites Getting started with Distributed Check- point (DCP).

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Getting started with Distributed Check- point (DCP)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.373829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.382598Z digest=sha256:e1b35adb7b6ec64a9e7622a6e51ff5e7b91275eeec689476d6a91b66580f8da9

Observation 5f164233-bc05-4186-942a-ec60c49025d7 · outbound

This paper cites Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.386378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.386378Z digest=sha256:3f7d0a72017d5f306d9e0d03dc8df2e1a2559ca6088e08773c94561930dc260c

Observation 9cfffabc-cb58-4e9a-b75f-bc7b8d852320 · outbound

This paper cites Simple is better: Multiplication may be all you need for llm request scheduling.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Simple is better: Multiplication may be all you need for llm request scheduling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.359485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.390858Z digest=sha256:45ae8103a66e089d0fac1953970751eafd5b4ce89e1a13c7e789f8042de7b15a

Observation 8a62c2a8-f310-4c44-aae6-6933689baeb9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Efficient memory management for large language model serving with pagedattention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.347609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.394615Z digest=sha256:28147fabef31ccff489afda0703f9e7fbd45cb464fabb96ece286961ceed9f84

Observation 48fdc530-c428-411c-87e2-112cddc9bfd4 · outbound

This paper cites https://docs.vllm.ai/en/stable/desig n/prefix_caching/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.vllm.ai/en/stable/desig n/prefix_caching/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.334793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.398238Z digest=sha256:619b25e473174b69b41d4ca7bdc44e66540ddc1434dda26efcdbd692566d9646

Observation 2ec7bafd-e84c-4950-a29a-579d9f6908fa · outbound

This paper cites https://lmsys.org/blog/2024-01-17-sglang/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2024-01-17-sglang/

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.322790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.401764Z digest=sha256:1ee3e2839c43531480b554b8a988da079297e9bc45c43e81fca989b6188490d4

Observation aed7ed80-0a54-4021-9bad-50150fca6b34 · outbound

This paper cites Pie: A pro- grammable serving system for emerging llm applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Pie: A pro- grammable serving system for emerging llm applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.310698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.405393Z digest=sha256:4d5ee5f3a24bd11e9bd03351e062e021dbc7adb5ddc7ca528c6db151b020ce76

Observation 2d7fe69a-4765-4fa5-a54c-4bd2caefbe3d · outbound

This paper cites Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.409426Z digest=sha256:4c15e3cdfc107685b54f27aa0d2c88d03cffbb4f9c889a775c32d21e5f445b71

Observation 4c63f9c8-ebc7-45d6-aa2d-1fb0fd3500ba · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.412959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.412959Z digest=sha256:3fd9f5dabaed027f755680eb68bd8b5e6d52f3206f25df3fccf8202f74771019

Observation 60382737-0556-4911-ab3d-b66a51c7f2ec · outbound

This paper cites Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.282711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.416940Z digest=sha256:8d8f753ee97efa092b20284c66e2ee91155b83bdaccde89132bed1406ab38d14

Observation 4c583cb5-6425-44b7-82ee-ab34222e266e · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.421353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.421353Z digest=sha256:d043e657be1a686002c883f263f2315407cde77bcfef03085d84e7a0ab1ac8c9

Observation f8c235cd-4da4-4fe4-88dc-0cf04bcc8159 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.425755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.425755Z digest=sha256:532a0b6655591e91a6521a787458eeb85e3cb4975ca8106303b0f65fa2429ca0

Observation 2847fa79-e643-40e6-a65c-fb56839f6f7e · outbound

This paper cites Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.430349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.430349Z digest=sha256:2c60f300b26f2460fb8ee1b92384d629c300900b205d615decebdb95292eb8a7

Observation 50a88b26-30cb-4888-a5af-cec5627c30a6 · outbound

This paper cites Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.269755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.434044Z digest=sha256:fb96bc82fbf12fb906f91d20e718bdace550aef3f1f74242654ec510cd26104b

Observation 006bb2b3-6cfd-4972-844b-26a05685227f · outbound

This paper cites Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.438822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.438822Z digest=sha256:0fb8b95314066e9e8bf310c867b42b024ce6cc922a6f6867936cebb29a3d01ad

Observation 32b178f7-8443-4441-b96b-3e37cc30c25a · outbound

This paper cites In search of an understandable consensus algorithm.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In search of an understandable consensus algorithm

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.249413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.442644Z digest=sha256:6be3a3d3d69f4585682b8d849e6abacbac2b1e58332014bba6cb6b60f7ec8ffd

Observation bbf2b854-54d3-4aaa-bb1d-fe5f9036b578 · outbound

This paper cites Chain replication for sup- porting high throughput and availability.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain replication for sup- porting high throughput and availability

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.238009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.446537Z digest=sha256:41c887f8b7f8e5fd79986853ff2234450b57661fa6538ec7753b7d26fe0af651

Observation ff7f2fae-e973-42ba-8053-ca9eff8b8ff9 · outbound

This paper cites Object storage on craq: High- throughput chain replication for read-mostly workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Object storage on craq: High- throughput chain replication for read-mostly workloads

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.224973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.450954Z digest=sha256:3130455c98568cc29b37d0e3ccf1d08586f97edaa86e1356a603f1bf284028d2

Observation 9168acd8-ed5d-4b95-8f5f-2019e44874c8 · outbound

This paper cites https://duckdb.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://duckdb.org/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.211272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.454631Z digest=sha256:92ccf8b53c2091d69c26587aaf18ab4c18278db5d5704c1abeb1d4a50d7c2a9f

Observation 53bf552d-21dd-4d45-a1ec-41026a92b6c8 · outbound

This paper cites mtcp: a highly scalable user-level tcp stack for multicore systems.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure mtcp: a highly scalable user-level tcp stack for multicore systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.198701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.459162Z digest=sha256:3733370f43b61497b1dbd1df2b31b2d509e794b4f78e49d94719939a727037d0

Observation 6a330889-e78c-40b5-9722-42ead742e5d1 · outbound

This paper cites https://github.com/juicedata/juicefs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/juicedata/juicefs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.186247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.462909Z digest=sha256:ccab26bdd7f7c32cb4cc137432fa22fe562d04bdc7ed4a5d16da742b5a94c597

Observation e5e4bb20-5e0d-4601-ba22-8341ef673049 · outbound

This paper cites https://github.com/scitix/InstantTensor.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/scitix/InstantTensor

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.172944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.466659Z digest=sha256:9c58cdffffd63f17816c7bbe3b4d8395f3f1b109335989e79c619eab8f314077

Observation 3e039b24-fd2d-4026-b25b-8d5d812f2fed · outbound

This paper cites Long- bench: A bilingual, multitask benchmark for long context understand- ing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Long- bench: A bilingual, multitask benchmark for long context understand- ing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.160803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.470342Z digest=sha256:3825497c77b5a886473aa6e944a67dd9db46ace7e0e7a5f0c1f8c853d111a978

Observation 39ad49e7-c9ab-40df-ac0f-262f08e9d101 · outbound

This paper cites Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.474373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.474373Z digest=sha256:3a554bb51048f6bf0d37029a4ad2acec84d280288bb3ae9c647508fd74a35929

Observation 22a3cb25-dbb3-44a9-83a9-ba3209d9e25d · outbound

This paper cites https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.147703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.478209Z digest=sha256:4e791aa14e48bd2d46f3e90d6aa01de4a6d22f956418dfd38d7949fd1bee732a

Observation f82df5cf-a17a-497c-a3eb-179e3270f03c · outbound

This paper cites The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.135851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.481823Z digest=sha256:5b621ae60c902f4518c893716aa0d5fcdfe0b51cf4f8f43a0a5a0b2e759d0799

Observation 916f0139-d982-41e4-9dda-0c918187b03b · outbound

This paper cites https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.123576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.486166Z digest=sha256:743ee121adf8350176c7d0f0b0b1b49e68c8586feec9b9dc6de377b407f6b91c

Observation ad3735ae-f21e-46a5-9c1d-c929ea663c71 · outbound

This paper cites Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.109945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.489786Z digest=sha256:f7cbfa4d710aea5edd2c2d02b08f34833c3f68e2f3a2377287da2e0101ca386d

Observation 50970dc1-da77-4f6d-87a2-d011f79d8682 · outbound

This paper cites A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.092229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.493268Z digest=sha256:910519c2b4a3457e14a73b0358a3849bfb440833cacb706191869ba39bfa6cdf

Observation 9da05baf-c805-45e5-b197-927e82cb80a7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.497271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.497271Z digest=sha256:62cb4bde996cc1be02aaf32def951238a2fb0776c5c5fa659836e813473ba848

Observation 7c371ba0-aef4-4f2f-9ba9-cc684d7c79ce · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.501915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.501915Z digest=sha256:022b758e3da9f91f0521c6b1acdb78192ce3b8daac955acfe9e8da2aef5a7021

Observation 3f1f042e-f93e-4b13-97de-451d75f5d386 · outbound

This paper cites Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.070297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.506285Z digest=sha256:838f7528ca7327e7955a61a51976db1e5f46d53b3ab2409faec0bc875201789a

Observation ec3db6c5-8618-41dd-a0e0-fc2975811352 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.056029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.510490Z digest=sha256:0ad0d38a1e8951e2b45ed73fe2f3c691c3333b7a4c91d008e2c84886fd533cfa

Observation ce5ad13f-c44f-4996-b5d5-3d1c1068e11c · outbound

This paper cites Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.041553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.514615Z digest=sha256:0f9a2be222d35a9fffd9c32662678a82f24e29d8f3608cc3704263eb29e0ded0

Observation 321a3a82-8a7f-431a-b288-9d475621f8dd · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Fast Distributed Inference Serving for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.520249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.520249Z digest=sha256:fcfa7fee6deef7b3b5dcc5be68ffbb199485de441aefc3211cb45d11897408e8

Observation d44d2940-594a-4da8-9e92-b240f4b8893d · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based}generative models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Orca: A distributed serving system for {Transformer-Based}generative models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.028543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.524500Z digest=sha256:977167830f62f438a668c459f3726286ead3b1ff7cb884afe29c5b1ce8b774cb

Observation e5b698e6-323f-4355-87d5-cc3ac62d84e0 · outbound

This paper cites Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.016099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.528420Z digest=sha256:85cba6266c96a733af10e0005c497cb1fc338ea7d48edcb48b608097f9f295ea

Observation a47f28ee-3dff-4fe0-b73d-64c8f230a1b1 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.999320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.532162Z digest=sha256:0d7436265478b80cf3820ef8959acde492adc4e0b203ad12aea6ebab6961672f

Observation d7f6d99a-bfde-4c43-b0d1-93c0571d3a74 · outbound

This paper cites Jenga: Effective memory management for serving llm with heterogeneity.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Jenga: Effective memory management for serving llm with heterogeneity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.985186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.535995Z digest=sha256:9315250e907acb5fb8f38f97812cf4b1e4081f401199a44de80a18cbcaeec461

Observation 47d81dd6-1a6f-458a-91d8-0c0eb459dba6 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.539810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.539810Z digest=sha256:dab891b3c3b109aa1ceaf1ebe14f5568aaf4d31c34da1888d4c32ef0166870a9

Observation 1ec8b77c-2dc6-4430-a65d-ae0ace596689 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.544105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.544105Z digest=sha256:ba1ae3e6cf42914ac89b07dbab25d6dea6b181c9fa78581fd5fa35502edb7301

Observation e727258c-6057-4d0d-b6fc-f4007d5afa6c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.965148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.547722Z digest=sha256:6e71fe31d6709f22677a2216035e38238717e09b99f2e891b6891e896dcb1f58

Observation 84c4db5f-97ed-4335-a15a-5f09fc966fd5 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.551553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.551553Z digest=sha256:66b6c983d4984b8a594605101956937499bbd37cea109691717d5a889b52dd0f

Observation 097379e1-7500-43d0-8074-29ede2225e43 · outbound

This paper cites https://gith ub.com/langchain-ai/langchain.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://gith ub.com/langchain-ai/langchain

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.953397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.555354Z digest=sha256:a12855de14d3f64cc37a4f4e29c4088bab6af9ffac2a55e59badb0204465f3db

Observation da08e753-aeb8-4041-b346-72bca9005c74 · outbound

This paper cites https://www.langflow.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.langflow.org/

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.940054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.558727Z digest=sha256:990eb55abd9de55fe894a7f063b86a2f1a4bc1099ca8741994d28175af5b7b36

Observation 829b080f-c9c2-465c-af2a-a92e10744480 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.562330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.562330Z digest=sha256:e9c58a0b21836292bff8c614a1dd71794af4e241fee09f58cc05494741f5d706

Observation a982b658-0abc-40b2-bc57-01279d88e8c3 · outbound

This paper cites Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.928405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.566354Z digest=sha256:4f248328d35d628a204ebd6a19998b198658f84f6c8d95d412a43b4d9747ca18

Observation d6368581-12ab-416e-8120-5c3d69cb5437 · outbound

This paper cites A System for Microserving of LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A System for Microserving of LLMs

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T18:23:08.608861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.569863Z digest=sha256:993e015245fd0538c8bf3c4497abfacb71f8ab2da4550998e1136bf57ba97a1b

Pith citing papers

No inbound Pith citation observations are available.