Pith. sign in

Paper Citation Record · LEDGER

Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

As of 18 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2403.08845.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.08845 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T00:38:49.320645Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:05:04.535360Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c8c871b-8641-46e0-85a3-6b05dc617ffa · inbound

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling cites this paper.

Large Language Monkeys: Scaling Inference Compute with Repeated Sampling Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:42:23.420509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:42:23.297389Z digest=sha256:3367497a5474ff9ebbafdfad5a3747a703598091475227c6087908ac5c8b3d14

Observation 17988e46-c51e-4b4a-a168-af462580bb22 · inbound

A Survey on Large Language Model Acceleration based on KV Cache Management cites this paper.

A Survey on Large Language Model Acceleration based on KV Cache Management Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 247

Resolution
unresolved
no resolver link, observed 2026-08-11T00:38:49.320645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T00:38:49.320645Z digest=sha256:b5889cd7c56316607392a6e3ea08fab29fb9ef940700091628bc1318033543d5

Observation 23078cb0-277e-4b00-ae7b-d36a4598bd3d · inbound

A Survey of Context Engineering for Large Language Models cites this paper.

A Survey of Context Engineering for Large Language Models Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:58:45.483712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T20:58:45.060041Z digest=sha256:4aad5b19a6a0fdfa11cf98044f506b64a8843a725bbdcb284923b846c1582dfc

Observation 9b20d94a-414c-4ffe-b039-789b68f017aa · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-19T15:52:38.676180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-19T15:48:59.800756Z digest=sha256:742150b2ea5c428315a9f27f1d6a7ff3193b4739b36eac3b531577928792c2f6

Observation ad8cbe96-ae1f-49ef-911d-2c1c23c10c7d · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:05:04.537324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-30T20:58:51.141871Z digest=sha256:a47d1f87590754b160268d8a59197ba6730ddd3bb1c92289ac9949c21cd91514

Observation 64b4fc2b-ae41-42ca-bbfb-24bd1691324a · inbound

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts cites this paper.

DualKV: Shared-Prompt Flash Attention for Efficient RL Training with Large Rollouts and Long Contexts Bifurcated Attention: Accelerating Massively Parallel Decoding with Shared Prefixes in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T14:03:08.545101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:03:08.545101Z digest=sha256:6ac222c8b6b43441bfd89cafa15ffcc52da0f338b03faf72eb3c105b6195362c