Pith. sign in

Paper Citation Record · LEDGER

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention

As of 5 August 2026, this Paper Citation Record lists 5 of 5 outbound references and 1 inbound Pith citation observation for arXiv:2604.07815.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.07815 v1

Coverage vector

measured 5 of 5 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T17:57:59.636975Z

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:03:08.617257Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-30T07:04:20.894665Z

Reference resolution

5 of 5 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch5

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 873da5d2-89b0-4e91-b6bb-4460768145a0 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:45:57.795459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:a68ebe0973b53f88e1638e23f8451e2493535631a56bd7f35a3dbd0f79fb50a4

Observation 19feda74-160d-46fd-931e-f6aa7ca968ee · outbound

This paper cites RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention RepoBench: Benchmarking Repository-Level Code Auto-Completion Systems

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T22:29:52.776747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:6e542d2eeb2af12376a58e1b5a16b2288b934471ae120144bffaf2c2013e03c9

Observation f1237d17-208a-4273-b25c-402907d5134f · outbound

This paper cites Longcat-flash technical report.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention Longcat-flash technical report

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:45:57.774641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:ed9fbeeb345f32399687246a36111e9d0e5689a68a6ce664fb271408f4ea3f86

Observation f7f8442b-e6d9-4ea7-9d6c-ac4d4c74212a · outbound

This paper cites FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPU

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:45:57.821474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:65421378a0e8d80647e3e24dc87aea783fe968597a0c1acd5d94c8de922941fb

Observation 000bc7c1-7a39-4beb-a902-2864270f8062 · outbound

This paper cites Qwen3 Technical Report.

AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention Qwen3 Technical Report

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T05:45:57.811820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T17:57:59.636975Z digest=sha256:69e2368ae981d40a7a04409756529fe9f73cd40eb2d0af6013843c80370186e1

Pith citing papers

Observation 9b91a285-c214-4683-987c-f4a163d3c9db · inbound

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding cites this paper.

Predict, Reuse, and Repair: Accelerating Dynamic Sparse Attention for Long-Context LLM Decoding AsyncTLS: Efficient Generative LLM Inference with Asynchronous Two-level Sparse Attention

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T07:04:20.896292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-30T07:03:08.617257Z digest=sha256:a15930b50966619b939345a0c31dd8c8f8a2290a4fb0bc1880e841194da8d387