Pith. sign in

Paper Citation Record · LEDGER

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication

As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2412.20501.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20501 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:26:40.947047Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da331029-1b30-41aa-aeb3-45a1bcdb499c · outbound

This paper cites SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.885542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.885542Z digest=sha256:42f8a549f0727b14c3085f3f23fd8eabe8b0bdd3bb322d5612a37e758035a00c

Observation a5f4dc18-95ea-4a28-9974-68f6d893dfd0 · outbound

This paper cites Striped Attention: Faster Ring Attention for Causal Transformers.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Striped Attention: Faster Ring Attention for Causal Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.892673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.892673Z digest=sha256:e172f1cbc542f9bd902e57aa2f66118b70df53184e331f40aae49605619d20fd

Observation 7d92dbd6-302e-4cc5-ae6d-11bc11f9ff21 · outbound

This paper cites D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.896082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.896082Z digest=sha256:3070ce66bfdb7b0c76d93e84107f1ee412e4580afc3211d5282eb358e54e32b0

Observation 6036e558-6c65-4d71-b8d5-6af65e7ad13a · outbound

This paper cites xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.909528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.909528Z digest=sha256:3aa1c0ab0b5659d67343830564a9c6b02b6cfbb75bda26e2de3c5eb7579c107a

Observation a16b346b-bd2d-496b-862d-91d13444cc93 · outbound

This paper cites LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.913166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.913166Z digest=sha256:614fd2ce61645fab3dd714992dcc17cb23bc2de32f15b9ab6fdb6cf9f0125c76

Observation 49cb6187-7629-4237-b25b-f64a9d094499 · outbound

This paper cites Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.923614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.923614Z digest=sha256:94ac7c75f1410de51e14ad4e29215291148cdc57e822c833eb47bbc369644006

Observation 7913a40f-7474-4fcc-8851-4e5844d968f8 · outbound

This paper cites Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.930670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.930670Z digest=sha256:0e24b1bdda302f18a21451ba4588169e5af1481c1e8826d0e43e66e8cbfd188c

Observation 2eef1bec-a9f9-4916-97bd-5e3d8bae575a · outbound

This paper cites Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.934049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.934049Z digest=sha256:3e83ed9039d2aaeb1ec6199f07d00e7abab8a6365ee82156ded75c6b5e069c15

Observation 3bfd4ff6-2d09-44e7-b018-90d4fc4d2fc6 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.937348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.937348Z digest=sha256:423408eb899a3c589896a6ec67feca1b6784a1f991a461dcc40d0a9afd491b76

Observation f5b1654a-de27-4a0f-9265-64732f89b371 · outbound

This paper cites DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.943865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.943865Z digest=sha256:da458d74990b669a8107220faa2d79a91ba2547740cbded277c5aef3c12cbf2a

Observation b1ea6812-da04-473a-8205-64a3fb3c26c6 · outbound

This paper cites You can have an appendix here.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication You can have an appendix here

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T23:26:41.146603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-10T23:26:40.947047Z digest=sha256:f93ad6b98ff5b894355eb1f3f400b4bc8c460e7ff243add8d3c68a83cdfa3cd5

Observation 39a86d49-20b0-49d8-8368-0e9abb0648fe · outbound

This paper cites The Llama 3 Herd of Models.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication The Llama 3 Herd of Models

Reference 2012

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.902785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.902785Z digest=sha256:cfdb267bafa68938c07aa6b018e9743146de9131ffa93fb8ba2638f5b24eeb77

Observation 8458eaa1-9adc-4e5f-8467-ae62d3015b5d · outbound

This paper cites Sequence Parallelism: Long Sequence Training from System Perspective.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Sequence Parallelism: Long Sequence Training from System Perspective

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.920005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.920005Z digest=sha256:13185eeeedea38bf3cb33314bfd310466e81bfe805a5ebeb0d0aee7ba0b34a68

Observation a2ef2535-1694-474a-893d-13668dfb5593 · outbound

This paper cites Context Parallelism for Scalable Million-Token Inference.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Context Parallelism for Scalable Million-Token Inference

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.940418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.940418Z digest=sha256:a725ce4e263d8b7ef33af29749b514903cfdcceae8d254e31f140d427dbdd0f8

Observation c264a08b-d356-4150-9852-c9575d148781 · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.916583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.916583Z digest=sha256:bc7ef7cc91ed06f308df04e585605c9872e4d4538067c817f04b5e314e952170

Observation fa18bbc3-f7d0-45c5-afa0-a46305ab6548 · outbound

This paper cites InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.899331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.899331Z digest=sha256:17ac616c3d6cf2dc444236c4da075e1c62fbe8bc52e86f090b9de9c7f2925f9a

Observation febc9a7e-5730-45f4-81f1-e45cc476bc97 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Fast Transformer Decoding: One Write-Head is All You Need

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.927027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.927027Z digest=sha256:807a6d7f78ee81484b33dfc9ac32603401a0055aa71739486321fc0da21a3e78

Observation 3079faf4-e985-4460-8757-2575b41a88bd · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.889229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.889229Z digest=sha256:ec23d007137e5cbd274645bb7fcfd632df1a4715ac63c6de851a08253a594705

Observation c3eea437-b1cd-433d-a163-85554eb40f67 · outbound

This paper cites USP: A Unified Sequence Parallelism Approach for Long Context Generative AI.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication USP: A Unified Sequence Parallelism Approach for Long Context Generative AI

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.905999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.905999Z digest=sha256:3bd5b55a307c6293b8a5de5953c9b48ceb84e3d4252d053586ab7e325664b54b

Pith citing papers

No inbound Pith citation observations are available.