Pith. sign in

Paper Citation Record · LEDGER

Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2408.04093.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2408.04093 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T12:36:41.329878Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T18:40:03.291994Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 9aba7698-778a-497d-99a9-9a11d8d01e13 · inbound

Star Attention: Efficient LLM Inference over Long Sequences cites this paper.

Star Attention: Efficient LLM Inference over Long Sequences Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T12:36:41.329878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T12:36:41.329878Z digest=sha256:14ea74bae612c4cdc60259e255f252f7cdbcbd3da80719c6f1773af03e902512

Observation 28e50240-751b-4c8e-ab2a-69bb2bc42f6d · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 257

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.513168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.513168Z digest=sha256:e970beb12258d2ca2089684d28613cbb884477556f2045500e48320b264303d1

Observation 7913a40f-7474-4fcc-8851-4e5844d968f8 · inbound

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication cites this paper.

TokenRing: An Efficient Parallelism Framework for Infinite-Context LLMs via Bidirectional Communication Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T23:26:40.930670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T23:26:40.930670Z digest=sha256:23faf253dac0d5c36b31bffe451a0ee50c5bb5fae37054f0d719837498f62b07

Observation 45e330fa-b746-4e55-8786-03d9f1dd90ed · inbound

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference cites this paper.

BlockDialect: Block-wise Fine-grained Mixed Format Quantization for Energy-Efficient LLM Inference Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T22:44:30.161918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:44:30.161918Z digest=sha256:1a8fd84674f2b255f3ae94371b488f6e712f3523c09c4f061d24ee11ded05c0c

Observation 516599cc-a18b-4080-ae52-665b23f30ef4 · inbound

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? cites this paper.

GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity? Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T20:25:16.933498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T20:25:16.933498Z digest=sha256:e2d5ff0722c5d9dc33a89d17dffc069687bf3d0af75b6945c5bbf0f36c06e24a

Observation 447711e1-fe9b-42e6-9541-dab3815cd531 · inbound

NoLoCo: No-all-reduce Low Communication Training Method for Large Models cites this paper.

NoLoCo: No-all-reduce Low Communication Training Method for Large Models Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T04:20:03.307804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:20:03.307804Z digest=sha256:ab5532dd8578991808d96ee4daf4146781c1e949362859b53acbbe5d2518ca84

Observation f2b57c26-90b5-4e74-8058-7f15600fd2a8 · inbound

Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit cites this paper.

Pixel-Resolved Long-Context Learning for Turbulence at Exascale: Resolving Small-scale Eddies Toward the Viscous Limit Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:09:23.105792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:09:23.105792Z digest=sha256:41ac4aabdbe74417e18759b8d0c48b190b390be823fe1f5768dfb9a110a00ca2

Observation 4e204637-329e-483b-a9c9-18511a458a6a · inbound

xGR: Efficient Generative Recommendation Serving at Scale cites this paper.

xGR: Efficient Generative Recommendation Serving at Scale Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:54:17.726744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:54:17.726744Z digest=sha256:481dda5edf5a8e0878375bf8db72c7b6ac5871ae8ee4ac45f242939165567568

Observation b4c65c12-8ddc-4abb-bafd-2eec43b86923 · inbound

ZAYA1-8B Technical Report cites this paper.

ZAYA1-8B Technical Report Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 149

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:26:05.203433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-08T17:36:37.182196Z digest=sha256:da82b0fb8234ff91f8933a517abd63ffc9afec31d1264ac8639d4302f0ba1906

Observation dc3a5342-c229-4ce5-888f-ed62b265a4b7 · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-07-04T18:40:03.293373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T22:37:15.072758Z digest=sha256:751ae7abe5fa182b7d42deedf2ee897474d7a7a3e33c584b52a32c42a88c4329

Observation 123606a0-0f96-49fe-a6dc-9a75af427b0d · inbound

ZONOS2 Technical Report cites this paper.

ZONOS2 Technical Report Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:15:58.993906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-29T02:07:31.791835Z digest=sha256:2cdb3b14f6d6de3270cf727fec47b5919ff52e6a94066bf52fadd879618f6ae6

Observation 81252ced-3e2e-403a-93d2-58ccef9a2adc · inbound

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution cites this paper.

ZUNA1.1: A more flexible EEG foundation model for Denoising and Super-resolution Tree Attention: Topology-aware Decoding for Long-Context Attention on GPU clusters

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-01T09:52:04.811661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T09:52:04.811661Z digest=sha256:d581c26b50850d8ce65210f5bb59febca1a848b45ccf10fa33f3d748e03a788f