Pith. sign in

Paper Citation Record · LEDGER

Context Parallelism for Scalable Million-Token Inference

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2411.01783.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.01783 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:57:08.267655Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T07:27:44.457158Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 66ae41e1-028a-4b2f-a0b6-98292fad1375 · inbound

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training cites this paper.

Hardware Scaling Trends and Diminishing Returns in Large-Scale Distributed Training Context Parallelism for Scalable Million-Token Inference

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T16:57:08.267655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:57:08.267655Z digest=sha256:d749aa8d44bcd8bfe829ebd9cc199769f170fc4b5257f09f1684da10bdf6f8f8

Observation a2ef2535-1694-474a-893d-13668dfb5593 · inbound

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication cites this paper.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Context Parallelism for Scalable Million-Token Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.940418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.940418Z digest=sha256:e4132376b4886f2efc04dab571192452136752b49d07daf62802559c2212ff1d

Observation de723323-0694-46f5-8a4c-dcd90f7cf9bd · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Context Parallelism for Scalable Million-Token Inference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:59.378063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:59.378063Z digest=sha256:b87637fbe6b24b36570b05b8ba6eccd70c77fefdd582e8a9e7cf550d6461dd23

Observation a531bbb9-9266-4ee2-8fbd-2c505209c3fb · inbound

PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training cites this paper.

PRISM: Probabilistic Runtime Insights and Scalable Performance Modeling for Large-Scale Distributed Training Context Parallelism for Scalable Million-Token Inference

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T06:30:59.726041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T06:27:40.254366Z digest=sha256:d31bc7351aeb874dffc68bcc022d96890ee0a4781dd6faf40b912cea030eaa3a

Observation 855e07e4-b108-4c24-86cd-a60d72f0414a · inbound

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap cites this paper.

Design Space Exploration of DMA based Finer-Grain Compute Communication Overlap Context Parallelism for Scalable Million-Token Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T17:17:34.555509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:17:34.555509Z digest=sha256:0093105d6b4a50f7f3995733a7c9c599f6be12cf5d8bb45d220d2b35d7553923

Observation 04656ae5-6a52-436c-bd45-ecf99efbe174 · inbound

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models cites this paper.

Neural Attention Search Linear: Towards Adaptive Token-Level Hybrid Attention Models Context Parallelism for Scalable Million-Token Inference

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T04:55:43.758571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:55:43.758571Z digest=sha256:037f8d6b1771aa2ac1d069e03c1f46a9491783a834359b0d17738e46ccc66f6a

Observation c00ec107-823a-42a1-989a-50fcfe9ee556 · inbound

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction cites this paper.

Scal3R: Scalable Test-Time Training for Large-Scale 3D Reconstruction Context Parallelism for Scalable Million-Token Inference

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:50:59.717563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:56:14.380078Z digest=sha256:d217a827a4b7d10eac6b08ecda0e71133d6d91a8ae7e68866d8e9ba24ac28c6c

Observation 3bbf67b1-a67e-47ae-ba30-5b67cafbbf4c · inbound

Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving cites this paper.

Accuracy Is Speed: Towards Long-Context-Aware Routing for Distributed LLM Serving Context Parallelism for Scalable Million-Token Inference

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T08:17:37.475817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T08:15:45.474618Z digest=sha256:35dd2fb7c63d0117ab87d92425e23382007139d6f544f15a519ce333519ed3d3

Observation 5af68555-97f5-4986-ad77-b6ede239c494 · inbound

GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving cites this paper.

GhostServe: A Lightweight Checkpointing System in the Shadow for Fault-Tolerant LLM Serving Context Parallelism for Scalable Million-Token Inference

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:38:23.124324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T00:37:08.671539Z digest=sha256:9ba0d34094807d0e59641ab899e30f181fcbe2c43c93b2996dccd7ef93428972

Observation 892ee340-59b7-4735-b4d8-5966013d2a57 · inbound

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics cites this paper.

Move the Query, Not the Cache: Characterizing Cross-Instance Latent Attention Redistribution Across GPU Fabrics Context Parallelism for Scalable Million-Token Inference

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:56:15.703053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T16:02:53.294235Z digest=sha256:c52f44b5093bc75e47de30d40c9f5dc2f597e13cdbab5522349f0933043a0f5e

Observation 4b8db0c8-f2d6-4491-91e5-8e39f5b70d46 · inbound

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design cites this paper.

Achieving Cloud-Grade SLOs for Local Mixture-of-Experts Inference through CPU-GPU Hybrid Design Context Parallelism for Scalable Million-Token Inference

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-03T07:27:44.458605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T12:06:46.806138Z digest=sha256:8eb3290ed9c39ecc2496f0ced1209588ebb7f23efb5957b6a64521d81a051e08

Observation 3fc6e178-fbbf-49ca-9ec6-df41b844ad57 · inbound

Towards Distributed Inference of LLMs on a P2P Network cites this paper.

Towards Distributed Inference of LLMs on a P2P Network Context Parallelism for Scalable Million-Token Inference

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:15:07.945392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T23:13:57.705287Z digest=sha256:5a2c4b15b496e19db1acfb70a93190be81487b1fb106ac27796d15cea1a482fa

Observation 744dfe47-89ce-4be2-8d8f-d136776e890b · inbound

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding cites this paper.

KernelFlume: Elastic Core-Attention Scaling for Agentic Long-Context Decoding Context Parallelism for Scalable Million-Token Inference

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T03:04:14.241549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T02:54:11.777615Z digest=sha256:a90d504f01cb320925003bdc717b5c8cadf3131959cb61f3c936ebf423e60e8b