Pith. sign in

Paper Citation Record · LEDGER

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure

As of 9 August 2026, this Paper Citation Record lists 88 of 88 outbound references and 0 inbound Pith citation observations for arXiv:2608.06007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.06007 v1

Coverage vector

measured 88 of 88 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T18:23:08.569863Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

88 of 88 outbound references displayed

  • verified exact2
  • verified fuzzy51
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7266694b-2167-460e-a9a2-c6a1461ddcae · outbound

This paper cites https://www.kimi.com/blog/kimi- k3.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.kimi.com/blog/kimi- k3

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.210099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.210099Z digest=sha256:f8064abeb8b3aaecd1d50ab9026ddc2ff101c001d3b44c79b9dd2755f3e2333d

Observation c5640c60-ddb8-46c0-820b-289b6a800bd5 · outbound

This paper cites ServerlessLLM: Low-Latency serverless inference for large language models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure ServerlessLLM: Low-Latency serverless inference for large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.215746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.215746Z digest=sha256:e4d65223c93707c4292528adadbd45bc2df4373c0f889b0e46548960400cd713

Observation d373c80b-505e-4110-9072-7ef34e468468 · outbound

This paper cites Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deep- flow: Serverless large language model serving at scale.arXiv e-prints, pages arXiv–2501, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.220987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.220987Z digest=sha256:01d74eff59b6bbe6ee57b0bb6d421d16ef6feaf3be5c124d44c8cf1a266c036b

Observation 059a9267-99f7-4671-b484-7768bf48ee96 · outbound

This paper cites Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Torpor: GPU-Enabled Serverless Computing for Low-Latency, Resource-Efficient Inference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.225722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.225722Z digest=sha256:138d945ec257696d2931b6f57d66040283d46b8e469604aa5d1fb246153bc471

Observation 7bb633c6-8d07-43e5-a398-f51152382d11 · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.230680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.230680Z digest=sha256:5925979e12a609659163502f56536f16ee3b6d53d19589c917f99731988599bf

Observation e4436a8a-eac5-4d58-a312-6f25afa6d915 · outbound

This paper cites BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BlitzScale: Fast and Live Large Model Autoscaling with O (1) Host Caching

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.234958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.234958Z digest=sha256:5d3819d473b5fa9c9d33b1aa9a28c914238bd1ae46c85f7a97be52fd8ce5ea69

Observation 4c447d41-9049-42d6-a526-314478418a4a · outbound

This paper cites Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Hydraserve: Minimizing cold start latency for serverless llm serving in public clouds.arXiv preprint arXiv:2502.15524, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.239382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.239382Z digest=sha256:042d571a5ab4a67b5e86a4fbd6a8c0966d8cdd5a1c46b1119851716f00dd99a6

Observation cff16d1b-78d6-4d53-afdf-1a465b2f75e0 · outbound

This paper cites https://lmsys.org/blog/2025-12-10-rfork/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-12-10-rfork/

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.243347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.243347Z digest=sha256:c7caff25bfd61cf1ed3d51ad56a434f5e1ea6738e445b59c8025af74d6d4bc59

Observation 9012b892-fc8d-46fe-9e9a-83c5ef0d592a · outbound

This paper cites Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Mooncake: Trad- ing more storage for less computation—a KVCache-centric architecture for serving LLM chatbot

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.247723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.247723Z digest=sha256:e5efc90e408c234ed08c72b18ce74cef9b33813772d38f14a908d7b37ddfc3f4

Observation 214672d2-ef54-4bbc-9f95-a8ecd3432886 · outbound

This paper cites Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dualmap: Enabling both cache affinity and load bal- ancing for distributed LLM serving

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.252135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.252135Z digest=sha256:20ad602f2970486d1e3fbcfc1e46feaf03d8447c98d9d8f0c5603c6787573c1f

Observation 530077e3-3828-4a3b-8d88-67308194f45b · outbound

This paper cites Lmcache: An efficient kv cache layer for enterprise-scale llm inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Lmcache: An efficient kv cache layer for enterprise-scale llm inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.256927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.256927Z digest=sha256:149a9ed743cb462390a3e95561a99eb9b37f70b737c72098ba65daf75a212726

Observation 1bab6d33-92c7-4dee-8e00-57df7e7504a8 · outbound

This paper cites https://lmsys.org/blog/2025-09-10-sglang-hicache/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2025-09-10-sglang-hicache/

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.261289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.261289Z digest=sha256:863cfcc0af8d1a8b09688a34db4a272c038169393a125cbfb72395f5c0bc6d5b

Observation 42f290ec-4062-48e3-bfca-971efb32b231 · outbound

This paper cites Stateful large language model serving with pensieve.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Stateful large language model serving with pensieve

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.647674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.265567Z digest=sha256:dc3fe173848d90e9ce5bee7698e60e1c65f205d795f6c263aad235a691956184

Observation 9194a7a9-2f73-4e9c-b28f-1f56071a2766 · outbound

This paper cites Cacheblend: Fast large language model serving for rag with cached knowledge fusion.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cacheblend: Fast large language model serving for rag with cached knowledge fusion

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.635789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.270101Z digest=sha256:526afecb611d25e4f62389db4329eab4f3ef64e6c5d1c54809ac6d8414a0ff3b

Observation da5d784e-d1e4-4906-82b2-e611454f3ffb · outbound

This paper cites Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient large language model serving for multi-turn conversations with Cache- dAttention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.624659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.274263Z digest=sha256:176fffebcb878fb223201b3cdfa16fe3dbcc36eea77f3baf97257188e772678a

Observation 7b20c028-7b9f-4073-9b77-231851ecc7da · outbound

This paper cites {ByteCheckpoint}: A unified checkpointing system for large foundation model development.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {ByteCheckpoint}: A unified checkpointing system for large foundation model development

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.613146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.278745Z digest=sha256:53229f31a2cfeb2d2df3c13fa48bc4144b6929e4e90b8411dc46251a4345dd1f

Observation 4662b785-a21c-4aa6-a755-bec13fc799eb · outbound

This paper cites https://github.com/MoonshotAI/chec kpoint-engine.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/MoonshotAI/chec kpoint-engine

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.601572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.282891Z digest=sha256:027d82f64bac8a0b0bac099ec8223381de5233c0b117e740ee848b05547f78c6

Observation 31c3bf0a-5caf-4897-8dfd-3bc84f726e13 · outbound

This paper cites https://vllm.ai/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://vllm.ai/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.589655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.286808Z digest=sha256:812c1dfdec7c199cb171eca48c0ddc66573fada8a10c736bfcacec46f554cb98

Observation 4559eff5-182a-4fde-b2db-6844d6b4eef1 · outbound

This paper cites Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Sglang: Efficient execution of structured language model programs.Advances in neural information processing systems, 37:62557–62583, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.292156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.292156Z digest=sha256:1fc92015b1fc676afa563ac93599c3a34fe197ef918316089f9dd63eb569c236

Observation f1a612b6-281d-44a8-966a-de5cd8b8260d · outbound

This paper cites https://nvidia.github.io/TensorRT-LLM/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://nvidia.github.io/TensorRT-LLM/

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.571245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.296384Z digest=sha256:fdecae8430b986b01fe016134759bb98d9f064f82a604b25d72ff9df55971506

Observation 88da8772-4852-4076-aeed-3d06e9d38c5c · outbound

This paper cites an unresolved cited work.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-07T18:23:10.559915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.300372Z digest=sha256:f9570d2cd081d16a92b65f178d3f4ba188aeb339d4dd28bdad9f63d9f882e18d

Observation db842daf-c122-4ea2-933e-7cb220a4e955 · outbound

This paper cites https://redis.io/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://redis.io/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.546071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.304346Z digest=sha256:3bfd4738185610fa5a9d86c8f3ef0c64072baf9cad2d240538975120aa6c37a0

Observation e67a7a63-1040-4cf5-a8c1-e12cedf1315d · outbound

This paper cites Vineyard: Optimizing data sharing in data- intensive analytics.Proc.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Vineyard: Optimizing data sharing in data- intensive analytics.Proc

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.534517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.308377Z digest=sha256:d346bca100ad7f9bea8571629f0a69ee30314293ec344692b443131e88565d49

Observation 45e262c5-30d2-4c6f-be79-3d1fc74ab155 · outbound

This paper cites https://github.com/ray-project/plasma.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/ray-project/plasma

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.522032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.312295Z digest=sha256:b212db5826dd17b55640262e4de8ffe2ad2ee8a796987b8263784f74ccfc1868

Observation 8436dcad-3029-4bdf-bb28-e3e864b50b22 · outbound

This paper cites Ray: A distributed framework for emerging ai applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ray: A distributed framework for emerging ai applications

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.509673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.315974Z digest=sha256:458ae9bdc51d40e3e738282c13a852f819ad66964f06733566d099e96aec1c21

Observation 76dc0904-c3ee-4e49-a3cb-3e5b4a919e09 · outbound

This paper cites Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Resilient distributed datasets: A Fault-Tolerant abstraction for In-Memory cluster computing

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.497933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.320121Z digest=sha256:e8703a9e61c00040588e446af88404245a722ad5050cc1e335f7f0272b43f2bb

Observation 0d236e8d-f09c-4621-b24c-a6a9d6cc73ed · outbound

This paper cites Serverless computing: Design, implementation, and performance.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Serverless computing: Design, implementation, and performance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.324667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.324667Z digest=sha256:5da8de2e514fae09046ba2f55c41245f315d306b33f5b199f6f3c51bc83df9a4

Observation b4725148-3cf4-48f3-bcf2-2e285e89f710 · outbound

This paper cites Infless: a native serverless system for low-latency, high-throughput inference.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Infless: a native serverless system for low-latency, high-throughput inference

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.478626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.328764Z digest=sha256:337f23c38a6b308d623699f30d1a1bbbec86bf83e80c77d307db9900ebad7d1f

Observation 61af6026-a968-4494-9880-29304696fc9d · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.332432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.332432Z digest=sha256:29c023ad05b2cf2f8e8022484fd1eb1fd61c2a57e5f1b568cf0ba8155fb86c75

Observation d04194b5-e4b6-438e-8cae-de4d804b89f0 · outbound

This paper cites Language models are few-shot learners.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Language models are few-shot learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.336296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.336296Z digest=sha256:b4074ab1e1e59f06d56624c7ec6ae191f94f13f9c58933b0d45c6bd64e8b48e8

Observation 8c2c0de8-44a0-435c-8316-63e889a005a8 · outbound

This paper cites Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flexkv: Flexible index offloading for memory-disaggregated key-value store.arXiv preprint arXiv:2512.16148, 2025

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-08-07T18:23:09.340739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.339917Z digest=sha256:658e146c553e649620078a0ee6c00eed31cc0a2fd5f4d13a217e97b3126012d5

Observation 770044c6-56af-42e6-a0b9-808c69d64444 · outbound

This paper cites In USENIX OSDI, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In USENIX OSDI, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.450965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.344903Z digest=sha256:1c38c20c395a65401b219e31d58987ef85cb9a585e06db840b4a3d60d0b3ddf0

Observation c0f22ded-c796-44de-a7ef-351c921e3fd7 · outbound

This paper cites Splitwise: Efficient genera- tive llm inference using phase splitting.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Splitwise: Efficient genera- tive llm inference using phase splitting

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.438704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.348995Z digest=sha256:18d1f4f07d16d41ad42daefd79411bd71bb98dd3381f6433a49274e749c2a721

Observation ea77d351-5772-4605-b59d-d2342eac07e3 · outbound

This paper cites Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Inference without Interference: Disaggregate LLM Inference for Mixed Downstream Workloads

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.353271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.353271Z digest=sha256:2a9f6a122bbc24088314808a35068f8034de948232beb1e8b499f718d3e4ccc2

Observation 60fe65b5-c7d9-437f-aa3f-c6cb4b316c79 · outbound

This paper cites D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure D\'ej\`aVu: KV-cache Streaming for Fast, Fault-tolerant Generative LLM Serving

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.357426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.357426Z digest=sha256:80e4078379693eb44d5c31ce9b95ba2d0b8bc26719ec30f28bf9a32d52c4db87

Observation e3d048ff-38d0-4f2e-9257-ef5544fc095f · outbound

This paper cites Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prompt cache: Modular attention reuse for low-latency inference.Proceedings of Machine Learning and Systems, 6:325–338, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.425673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.361591Z digest=sha256:dfa73303371ef0b74c6ce65df87e94b46e065da0b965fc9ee1b7ea81cee77d8e

Observation bcc27414-be25-4ed7-90ec-d706b5fad965 · outbound

This paper cites Cachegen: Kv cache compression and streaming for fast large language model serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cachegen: Kv cache compression and streaming for fast large language model serving

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.414206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.365396Z digest=sha256:fab9acb6b02f19398eb821d321c7e5a08beef1be5ac476bc57c33a6f950df701

Observation 1a8aac5a-b716-4e1a-b453-3e151a60595c · outbound

This paper cites Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Ragcache: Efficient knowledge caching for retrieval-augmented generation.ACM Transactions on Computer Sys- tems, 44(1):1–27, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.401403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.369036Z digest=sha256:6e924fdcc64227c297978a7f7827893d3f727f18ae5c3598e8bbeac46856e574

Observation 573cad29-4ab0-4ad7-ae10-34ce8d5c21d4 · outbound

This paper cites Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Cost-Efficient Large Language Model Serving for Multi-turn Conversations with CachedAttention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.373466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.373466Z digest=sha256:4afa838fa3ffa31051cc41a39df09d247a6e6c620c555d412db49f0e1cfbc108

Observation 254654b9-7c4a-46cd-95fe-2db36adfceca · outbound

This paper cites Dist checkpointing package.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dist checkpointing package

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.386334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.378063Z digest=sha256:cba71a191061d180cefeee913c6482c5cfd3d75b3b7a04ac4a2286aaf01142d2

Observation 69313a8f-f463-45a8-ac16-3a10cb5bdf78 · outbound

This paper cites Getting started with Distributed Check- point (DCP).

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Getting started with Distributed Check- point (DCP)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.373829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.382598Z digest=sha256:f3f6d7a331ac167881bc1ae7d3c9256587e1386d67b238868107cfcdf7d15e01

Observation 5f164233-bc05-4186-942a-ec60c49025d7 · outbound

This paper cites Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Universal Checkpointing: A Flexible and Efficient Distributed Checkpointing System for Large-Scale DNN Training with Reconfigurable Parallelis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.386378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.386378Z digest=sha256:7f5236ee99132bd33394b21345883f3324adef27b79beffe3915b4dc3eed65cf

Observation 9cfffabc-cb58-4e9a-b75f-bc7b8d852320 · outbound

This paper cites Simple is better: Multiplication may be all you need for llm request scheduling.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Simple is better: Multiplication may be all you need for llm request scheduling

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.359485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.390858Z digest=sha256:658f928fd6cb5fab913cc2d84bf185e3e367d731e535cb51717494f74e7fada2

Observation 8a62c2a8-f310-4c44-aae6-6933689baeb9 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Efficient memory management for large language model serving with pagedattention

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.347609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.394615Z digest=sha256:9d35bde52ffaed9548f16ecf6b3d24bfb8f8889b9e9e00bcd6f9930fbc4b6f00

Observation 48fdc530-c428-411c-87e2-112cddc9bfd4 · outbound

This paper cites https://docs.vllm.ai/en/stable/desig n/prefix_caching/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.vllm.ai/en/stable/desig n/prefix_caching/

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.334793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.398238Z digest=sha256:ddefc07a7ba5e95b88e5230d0f8cc9e226b1e406bea76da57e677f6f8b031c36

Observation 2ec7bafd-e84c-4950-a29a-579d9f6908fa · outbound

This paper cites https://lmsys.org/blog/2024-01-17-sglang/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://lmsys.org/blog/2024-01-17-sglang/

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.322790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.401764Z digest=sha256:4496331f1455ed0e6977d1f79d11228601ff9f17bbb6395664800e51bfbff76b

Observation aed7ed80-0a54-4021-9bad-50150fca6b34 · outbound

This paper cites Pie: A pro- grammable serving system for emerging llm applications.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Pie: A pro- grammable serving system for emerging llm applications

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.310698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.405393Z digest=sha256:270b43065d80bf58f1bbea161d72e6df8917e98ecbff8c9c94bed41972a9fdd4

Observation 2d7fe69a-4765-4fa5-a54c-4bd2caefbe3d · outbound

This paper cites Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain-of-thought prompt- ing elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.409426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.409426Z digest=sha256:0ba62f5053d0efa5c3478a4f603e89b1b7a7f2c461c9b8e9e91fbadcabdae906

Observation 4c63f9c8-ebc7-45d6-aa2d-1fb0fd3500ba · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Tree of thoughts: Deliberate problem solving with large language models.Advances in neural information processing systems, 36:11809–11822, 2023

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.412959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.412959Z digest=sha256:52544abe626cdcf07b654a5d2825848223df1e941fed92c99d8abd063ee16ec1

Observation 60382737-0556-4911-ab3d-b66a51c7f2ec · outbound

This paper cites Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Training language models to follow instruc- tions with human feedback.Advances in neural information processing systems, 35:27730–27744, 2022

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.282711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.416940Z digest=sha256:4b8d8f204660c4359bd2eea2cc1ce0ac6620b33a2822bba1bc1dcf14317bfdf1

Observation 4c583cb5-6425-44b7-82ee-ab34222e266e · outbound

This paper cites Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.421353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.421353Z digest=sha256:15eacecf1255c1222543d16c17a1c2d8e5a36ae7818ffb785381c73f0ab1827a

Observation f8c235cd-4da4-4fe4-88dc-0cf04bcc8159 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.425755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.425755Z digest=sha256:4051c9082b3277a4501deb4c617096c35d5d22fbb707f0d66bfd44b721113a02

Observation 2847fa79-e643-40e6-a65c-fb56839f6f7e · outbound

This paper cites Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Totrl: Unlock llm tree-of-thoughts reasoning potential through puzzles solving.arXiv preprint arXiv:2505.12717, 2025

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.430349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.430349Z digest=sha256:ea46939f7b75ca54bcd451e34c297ab2a89050088fe057154d52c9bc758b6b69

Observation 50a88b26-30cb-4888-a5af-cec5627c30a6 · outbound

This paper cites Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Using name-based map- pings to increase hit rates.IEEE/ACM Transactions on networking, 6(1):1–14, 2002

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.269755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.434044Z digest=sha256:af45471a05eeb76ca0881fbd34a65a89a7b534482f9ccdb71c0c2541ee977061

Observation 006bb2b3-6cfd-4972-844b-26a05685227f · outbound

This paper cites Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Paxos made simple.ACM SIGACT News (Distributed Computing Column) 32, 4 (Whole Number 121, December 2001), pages 51–58, 2001

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.438822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.438822Z digest=sha256:e5dee98a53e15aec9a712b61f399e7f896e6449b4541b4a8b10347c52dbe5037

Observation 32b178f7-8443-4441-b96b-3e37cc30c25a · outbound

This paper cites In search of an understandable consensus algorithm.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure In search of an understandable consensus algorithm

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.249413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.442644Z digest=sha256:d864dad6f0931a32bf7ce0e6a2703a781ea0d4d65760a8ce6c9f881367b8af83

Observation bbf2b854-54d3-4aaa-bb1d-fe5f9036b578 · outbound

This paper cites Chain replication for sup- porting high throughput and availability.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Chain replication for sup- porting high throughput and availability

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.238009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.446537Z digest=sha256:21aa2f38f2c1a566a06726987a6fd4b59098b041ccfca9a8769f1eff5217304e

Observation ff7f2fae-e973-42ba-8053-ca9eff8b8ff9 · outbound

This paper cites Object storage on craq: High- throughput chain replication for read-mostly workloads.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Object storage on craq: High- throughput chain replication for read-mostly workloads

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.224973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.450954Z digest=sha256:b2fbf5f1f4e1be83f117c11f659d69e3eb8e61a8492b618d23728b8bffaf847c

Observation 9168acd8-ed5d-4b95-8f5f-2019e44874c8 · outbound

This paper cites https://duckdb.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://duckdb.org/

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.211272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.454631Z digest=sha256:997e33c83e8ba8e7c2eeabe0312cd695ce0f41fb9e73cffac96f10e9e80100e9

Observation 53bf552d-21dd-4d45-a1ec-41026a92b6c8 · outbound

This paper cites mtcp: a highly scalable user-level tcp stack for multicore systems.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure mtcp: a highly scalable user-level tcp stack for multicore systems

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.198701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.459162Z digest=sha256:acf0a7e620900eb75d6907884f6312ab3a24115c0693ee9fbb96f94cd91cf217

Observation 6a330889-e78c-40b5-9722-42ead742e5d1 · outbound

This paper cites https://github.com/juicedata/juicefs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/juicedata/juicefs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.186247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.462909Z digest=sha256:5b6cbb13ff837bf167c1186f9378935c6816f9b00caeeb0ef20393a01d749c24

Observation e5e4bb20-5e0d-4601-ba22-8341ef673049 · outbound

This paper cites https://github.com/scitix/InstantTensor.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://github.com/scitix/InstantTensor

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.172944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.466659Z digest=sha256:4d7ec71251a74de291f2742c6a981c120f9a3f32f34e6086117adf86e1805d72

Observation 3e039b24-fd2d-4026-b25b-8d5d812f2fed · outbound

This paper cites Long- bench: A bilingual, multitask benchmark for long context understand- ing.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Long- bench: A bilingual, multitask benchmark for long context understand- ing

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.160803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.470342Z digest=sha256:5db8fe95a12ace3a6e97b05054783c820d18684a8aba099cb44625bbb4e4a24c

Observation 39ad49e7-c9ab-40df-ac0f-262f08e9d101 · outbound

This paper cites Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-Datacenter

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.474373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.474373Z digest=sha256:48cee519c14da310ad5ca0008140a5d142ae8f6b77079fc858b6ac964fef12c7

Observation 22a3cb25-dbb3-44a9-83a9-ba3209d9e25d · outbound

This paper cites https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://docs.sglang.io/docs/advanced_feature s/sgl_model_gateway

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.147703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.478209Z digest=sha256:425ca5f0c218dfbf79baeb9169b1cd630e4f7deddfce20af2562e8c608b534b6

Observation f82df5cf-a17a-497c-a3eb-179e3270f03c · outbound

This paper cites The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure The power of two choices in randomized load balancing.IEEE transactions on parallel and distributed systems, 12(10):1094–1104, 2002

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.135851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.481823Z digest=sha256:2b2e0cee4ce413bce301c07ce0cf5b0fcb66341a56dc4cb440b7f58c9185ace8

Observation 916f0139-d982-41e4-9dda-0c918187b03b · outbound

This paper cites https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://huggingf ace.co/datasets/SWE-Gym/OpenHands-Sampled-Trajectories

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.123576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.486166Z digest=sha256:85aa06cd6a75d14630ebf96f620f14962b77b43e8d16ee8facffda601068fa00

Observation ad3735ae-f21e-46a5-9c1d-c929ea663c71 · outbound

This paper cites Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Statistical analysis of a telephone call center: A queueing-science perspective.Journal of the American statistical association, 100(469):36–50, 2005

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.109945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.489786Z digest=sha256:7435f7d6d06631696cef1a8e24c1f0c146c9097dd5e979ddd6d9fcd6296f6a48

Observation 50970dc1-da77-4f6d-87a2-d011f79d8682 · outbound

This paper cites A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A poissonian explanation for heavy tails in e-mail communi- cation.Proceedings of the National Academy of Sciences, 105(47):18153– 18158, 2008

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.092229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.493268Z digest=sha256:d162c22cedfdff60c8fca8322cedccfcef1e9ea203bb20f97002cf19dd0e77ed

Observation 9da05baf-c805-45e5-b197-927e82cb80a7 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.497271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.497271Z digest=sha256:dc646245e6f56647c9d4f3cea84d5fd4fc5cb23b4352ad8e59e4de2b683cff30

Observation 7c371ba0-aef4-4f2f-9ba9-cc684d7c79ce · outbound

This paper cites Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed: System optimizations enable training deep learning models with over 100 billion parameters

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.501915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.501915Z digest=sha256:d9c78c1481b3e700a73b3d2c9f2fe79a6ac0d4605ba23cefe92642ac590aa6b1

Observation 3f1f042e-f93e-4b13-97de-451d75f5d386 · outbound

This paper cites Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Alpa: Automating inter-and{Intra-Operator} par- allelism for distributed deep learning

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.070297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.506285Z digest=sha256:e13055f8133608a870db311d29bc64053970df15cb857c93164edb03b7c9ab6b

Observation ec3db6c5-8618-41dd-a0e0-fc2975811352 · outbound

This paper cites Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Deepspeed-inference: enabling efficient inference of transformer models at unprecedented scale

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.056029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.510490Z digest=sha256:22761300a05ca394ccd8b712460fc06eee3bd5d89bf939fbe7445c1ffcda9cf7

Observation ce5ad13f-c44f-4996-b5d5-3d1c1068e11c · outbound

This paper cites Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Loongserve: Efficiently serving long-context large lan- guage models with elastic sequence parallelism

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.041553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.514615Z digest=sha256:d525ea0bc261c03a61bccdea192b863c2cf847762ea4b46d317f617e7b357585

Observation 321a3a82-8a7f-431a-b288-9d475621f8dd · outbound

This paper cites Fast Distributed Inference Serving for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Fast Distributed Inference Serving for Large Language Models

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.520249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.520249Z digest=sha256:cdb712f87486aa19b7bb6401c17e9e64b4d8c5d1c5af64affd1bf367db198d69

Observation d44d2940-594a-4da8-9e92-b240f4b8893d · outbound

This paper cites Orca: A distributed serving system for {Transformer-Based}generative models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Orca: A distributed serving system for {Transformer-Based}generative models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.028543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.524500Z digest=sha256:dd663d9a2d0edf8c7440c1afb38acbc20e677b0ff4b2dfacde32efe948db9a03

Observation e5b698e6-323f-4355-87d5-cc3ac62d84e0 · outbound

This paper cites Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Taming{Throughput-Latency} tradeoff in{LLM} inference with {Sarathi-Serve}

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:10.016099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.528420Z digest=sha256:4cd02de3120b30c57ca7e9061d741f2ca718cc86dd947306ff2301fd6f4fc6f4

Observation a47f28ee-3dff-4fe0-b73d-64c8f230a1b1 · outbound

This paper cites {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure {InfiniGen}: Efficient generative inference of large language mod- els with dynamic{KV}cache management

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.999320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.532162Z digest=sha256:ec716555f404e98e8f2c1d01bbb5ba52d070587e97f9006a095a1f16e93bc3b1

Observation d7f6d99a-bfde-4c43-b0d1-93c0571d3a74 · outbound

This paper cites Jenga: Effective memory management for serving llm with heterogeneity.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Jenga: Effective memory management for serving llm with heterogeneity

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.985186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.535995Z digest=sha256:a0216e69f6fd27f3db543e6e3493d9e7a30d2cebe6c9ac5333e381610aade08a

Observation 47d81dd6-1a6f-458a-91d8-0c0eb459dba6 · outbound

This paper cites BitNet: Scaling 1-bit Transformers for Large Language Models.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure BitNet: Scaling 1-bit Transformers for Large Language Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.539810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.539810Z digest=sha256:16816b64d8f96795294044f1c6eff607f8366c5fe257bf013f2418549ebd5870

Observation 1ec8b77c-2dc6-4430-a65d-ae0ace596689 · outbound

This paper cites Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Awq: Activation-aware weight quantization for on-device llm compression and acceleration.Proceedings of machine learning and systems, 6:87–100, 2024

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.544105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.544105Z digest=sha256:30936cd8ce3e75495f205bfae2f68a0f4d7d8b3875f804ce39a941330c7e3503

Observation e727258c-6057-4d0d-b6fc-f4007d5afa6c · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Flashattention: Fast and memory-efficient exact attention with io- awareness.Advances in neural information processing systems, 35:16344–16359, 2022

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.965148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.547722Z digest=sha256:0c8906baa33062ccdb987fba19e2fb5e0a08f811d5b2778fc6b0cae7403ee71e

Observation 84c4db5f-97ed-4335-a15a-5f09fc966fd5 · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.551553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.551553Z digest=sha256:f677c80c1e7dfe07a150b4f69b23aa2429a15845a2d25801c9b5360ed533b9fd

Observation 097379e1-7500-43d0-8074-29ede2225e43 · outbound

This paper cites https://gith ub.com/langchain-ai/langchain.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://gith ub.com/langchain-ai/langchain

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.953397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.555354Z digest=sha256:b8565d035b3a7e315746ca9822a6afbf2777d0a02c78963abf8c9422b8a02009

Observation da08e753-aeb8-4041-b346-72bca9005c74 · outbound

This paper cites https://www.langflow.org/.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure https://www.langflow.org/

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.940054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.558727Z digest=sha256:c3a5ad6b2db47688a3eda8d55d6566e6e9056ed6cf63d563812b930a8b242e58

Observation 829b080f-c9c2-465c-af2a-a92e10744480 · outbound

This paper cites AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T18:23:08.562330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T18:23:08.562330Z digest=sha256:dba08d0efb7922413d73167cb7a0da2f07cb28a363d16f1368234787fbe3f7f2

Observation a982b658-0abc-40b2-bc57-01279d88e8c3 · outbound

This paper cites Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure Dspy: Compiling declarative 17 language model calls into state-of-the-art pipelines

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T18:23:09.928405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.566354Z digest=sha256:2faa404a1bc2ec982d282fc51ef123e3a38e5319a031ec596f12f01f322dda3b

Observation d6368581-12ab-416e-8120-5c3d69cb5437 · outbound

This paper cites A System for Microserving of LLMs.

TensorCast: The Missing Tensor Management Layer in Large Language Model Infrastructure A System for Microserving of LLMs

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-08-07T18:23:08.608861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T18:23:08.569863Z digest=sha256:be08fa41391f832dfd5ea2b6556d1a55ce6d4d0cf4808bc007de070cf871f7b3

Pith citing papers

No inbound Pith citation observations are available.