Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:26:40.947047Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2412.20501.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:26:40.947047Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation da331029-1b30-41aa-aeb3-45a1bcdb499c · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication SARATHI: Efficient LLM Inference by Piggybacking Decodes with Chunked Prefills
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5f4dc18-95ea-4a28-9974-68f6d893dfd0 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Striped Attention: Faster Ring Attention for Causal Transformers
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d92dbd6-302e-4cc5-ae6d-11bc11f9ff21 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication D., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., et al
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6036e558-6c65-4d71-b8d5-6af65e7ad13a · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication xDiT: an Inference Engine for Diffusion Transformers (DiTs) with Massive Parallelism
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a16b346b-bd2d-496b-862d-91d13444cc93 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication LoongTrain: Efficient Training of Long-Sequence LLMs with Head-Context Parallelism
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49cb6187-7629-4237-b25b-f64a9d094499 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7913a40f-7474-4fcc-8851-4e5844d968f8 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eef1bec-a9f9-4916-97bd-5e3d8bae575a · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Unveiling Redundancy in Diffusion Transformers (DiTs): A Systematic Study
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bfd4ff6-2d09-44e7-b018-90d4fc4d2fc6 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication LLaMA: Open and Efficient Foundation Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5b1654a-de27-4a0f-9265-64732f89b371 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ea6812-da04-473a-8205-64a3fb3c26c6 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication You can have an appendix here
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 39a86d49-20b0-49d8-8368-0e9abb0648fe · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication The Llama 3 Herd of Models
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8458eaa1-9adc-4e5f-8467-ae62d3015b5d · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Sequence Parallelism: Long Sequence Training from System Perspective
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2ef2535-1694-474a-893d-13668dfb5593 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Context Parallelism for Scalable Million-Token Inference
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c264a08b-d356-4150-9852-c9575d148781 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa18bbc3-f7d0-45c5-afa0-a46305ab6548 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication InternEvo: Efficient Long-sequence Large Language Model Training via Hybrid Parallelism and Redundant Sharding
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation febc9a7e-5730-45f4-81f1-e45cc476bc97 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Fast Transformer Decoding: One Write-Head is All You Need
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3079faf4-e985-4460-8757-2575b41a88bd · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3eea437-b1cd-433d-a163-85554eb40f67 · outbound
TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication USP: A Unified Sequence Parallelism Approach for Long Context Generative AI
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.