Pith. sign in

Paper Citation Record · LEDGER

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention

As of 5 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2607.25291.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.25291 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:55:55.511467Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved68
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6473973-a10f-4e0a-936b-8028a4469081 · outbound

This paper cites 2024 , journal=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention 2024 , journal=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.003957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.003957Z digest=sha256:2642dccacebc660a1f3214578c89195cec5e0c5f717aa091b63f43f0a1bc5988

Observation 57b10020-c421-416a-b7fc-22318a361253 · outbound

This paper cites SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention SeerAttention: Learning Intrinsic Sparse Attention in Your LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.085046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.085046Z digest=sha256:7831b3c9d072dc33c5ea9db48c7ee0fcd3250d8ef03c91f87ea249ed49a02489

Observation fdcc76a3-7adf-4d1d-bc3a-d79ce20d39ec · outbound

This paper cites XAttention: Block Sparse Attention with Antidiagonal Scoring.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention XAttention: Block Sparse Attention with Antidiagonal Scoring

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.171852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.171852Z digest=sha256:6da27ff08d0ccca66c8b1e181c050d9063a527917d123f0cafd86a9a112171a4

Observation 5b93182a-78a7-444f-a9d4-cc2e587f5eaa · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.233049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.233049Z digest=sha256:0456b7f216ccd22b0c9d9953af2b1908a4d350ec2a74e24da8813ea350b35abb

Observation ba7fe839-b193-4471-a8cc-6705ae71df6c · outbound

This paper cites FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.301880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.301880Z digest=sha256:3a8efab24197eabcce0f79c741815f10ec902c9614c6eec9b234cf1f62eec45c

Observation 5a1b6b55-8db0-4bee-b50a-90fe5400569d · outbound

This paper cites arXiv preprint arXiv:2503.17407 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2503.17407 , year=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.413649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.413649Z digest=sha256:ca62ae2674ec7a6e3227f16cd20b87bfc5ff368f5f56bf25fd24e8613b0956a2

Observation 12acc1de-2857-4e70-beac-b9d2a4576739 · outbound

This paper cites A Survey on Transformer Context Extension: Approaches and Evaluation.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention A Survey on Transformer Context Extension: Approaches and Evaluation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.497701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.497701Z digest=sha256:81c80d717f9cef92666292323d41d8f83242e6cc6a64338b037aca2088d49fd3

Observation edb776b5-be67-4fe6-a1a9-5b6c0b60fb7e · outbound

This paper cites SWE-bench: Can Language Models Resolve Real-World GitHub Issues?.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention SWE-bench: Can Language Models Resolve Real-World GitHub Issues?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.633183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.633183Z digest=sha256:ac6a74bc3808b3dfa4f912f7e381a06a85d6538af405bf00ff8f2112403f9d66

Observation c77a66cd-3940-4284-b563-85501e99e149 · outbound

This paper cites Qwen3 Technical Report.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Qwen3 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.766469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.766469Z digest=sha256:08cc329b38550efcf355cd69a80076df1fa671dfb8b9723b059cde3814ed6e47

Observation bd45346f-5472-44fe-adb1-259727119d52 · outbound

This paper cites Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.880834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.880834Z digest=sha256:ac738332aaa4be9e0d7dceafb9396f9aaa60d568e35b24fb963b8dfed32d7bab

Observation 4cfa87f5-f21f-45a5-8e30-1ff8bdefe454 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:46.966133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:46.966133Z digest=sha256:d6f25bf209c75ea6c6ff9f9e748105d70b0850d634ac53a1f156def9fb28e04b

Observation 99dc8954-2bc6-4890-b367-eb0b7806cd9d · outbound

This paper cites IEEE transactions on knowledge and data engineering , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention IEEE transactions on knowledge and data engineering , volume=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.098123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.098123Z digest=sha256:e402b4b56c94ff29ca1403825e2e78970f20fc165c9b34e85d91995aa2e4b483

Observation 3bc2cc2f-e686-4e2f-8e31-2970e4db5c73 · outbound

This paper cites an unresolved cited work.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.202653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.202653Z digest=sha256:3b772d5fddd08023256968160165a31622810f2a2e137d1df9692ec790bea3aa

Observation 73aa13b8-b2ee-4221-bf90-e162ad4030f7 · outbound

This paper cites Longformer: The Long-Document Transformer.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Longformer: The Long-Document Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.329691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.329691Z digest=sha256:7d375a5c86a08c662fb2bc1224402aa0ca19a1468f6fd3689763abb6798e4f4f

Observation bda7806e-3f46-4b92-a9ac-182ee6c5fe41 · outbound

This paper cites Advances in neural information processing systems , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Advances in neural information processing systems , volume=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.480738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.480738Z digest=sha256:02c70c16cf4dbb3bd4aea5baa4630330890878f8fb5ea47f395733a145f3136e

Observation 40e78ac4-1092-4d7b-8755-6881fb3e07e1 · outbound

This paper cites LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention LongBench v2: Towards Deeper Understanding and Reasoning on Realistic Long-context Multitasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.646679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.646679Z digest=sha256:d7eb03f6932533b8d472cf8466cbe3d05c98a09f930f73357f3896452377ca38

Observation 7ea9e06b-6ef6-466d-98fd-33cc72dc3340 · outbound

This paper cites L ong B ench: A Bilingual, Multitask Benchmark for Long Context Understanding.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention L ong B ench: A Bilingual, Multitask Benchmark for Long Context Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.843533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.843533Z digest=sha256:a95e5190c79bad3b657959d81ddc67d887c5538349380dbc46835c82038d2242

Observation de651917-7dbf-4ac3-ae8c-a2aa712174d8 · outbound

This paper cites The Llama 3 Herd of Models.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention The Llama 3 Herd of Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:47.941934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:47.941934Z digest=sha256:71a0dca8731993bcb617a458e13ba6fad8756f4c018071a368b23699420b96d2

Observation 7887b898-9dbd-4a80-b42d-b00b546af26f · outbound

This paper cites DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention DuoAttention: Efficient Long-Context LLM Inference with Retrieval and Streaming Heads

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.085704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.085704Z digest=sha256:b58aaecb3eaa2c03cc2643d3e9eed144a4ec9c203a0596325602c0c6505b4130

Observation e198e754-b5c9-457f-a40e-ade83f9fbeed · outbound

This paper cites arXiv preprint arXiv:2406.14909 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2406.14909 , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.184773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.184773Z digest=sha256:0e004fac82d703ae30a16c8f375e4e4c06bf0d6c2b4ed96ce2dfc161d053f314

Observation 314b9551-7096-46a5-b542-c158c3b8d2d9 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.331046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.331046Z digest=sha256:83e81d70193811d79658ad46cab6d89231e248089f3581ee612b7c30bb80e05a

Observation e4fc1560-6eaa-4bb2-9fc6-c67a846795a2 · outbound

This paper cites and Ermon, Stefano and Rudra, Atri and R.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention and Ermon, Stefano and Rudra, Atri and R

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.470417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.470417Z digest=sha256:d7c224398b234b7bb774a349d1b0d5a519a47dc0d4b0a6d4bad14a32c4e70c46

Observation 6df14abd-0b68-4617-84e4-20a62a65e62b · outbound

This paper cites an unresolved cited work.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.556427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.556427Z digest=sha256:e4f6b2bab426e25dd2e632cdd9079794f9de71c9fa92c50fdf7f8b8da98e0b7e

Observation a0945939-3453-46e1-b22d-bbcaf1bf4fe7 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.670740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.670740Z digest=sha256:dfdb1f2942b3f7c8cfea71c02b61f7d51b70c5076d26078b4206453c9e9f27d8

Observation 0396ea82-1da0-4761-9476-e3525300648e · outbound

This paper cites arXiv preprint arXiv:2603.05451 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2603.05451 , year=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.807256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.807256Z digest=sha256:57a01869a874246c052f8debe49690b6cb7add7246a6784e8c6dc1ca20bb76ac

Observation 4eecc743-4bc8-41cf-b075-d1298e6c2905 · outbound

This paper cites Advances in neural information processing systems , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Advances in neural information processing systems , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:48.932050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:48.932050Z digest=sha256:b6f9929217d2220ebd506ea5865fdd2dfb1c470a453c74cdcc03198c3dcb39b9

Observation d5c12d74-e239-4c32-97d9-25f5c1bc1304 · outbound

This paper cites Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.035284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.035284Z digest=sha256:798c7b31c13e5d4d30cf7ded18d89d313672e93d1ff75819ff794b6951f4400d

Observation 1ce58eaf-ac3a-466a-b5f7-694ebba9ac76 · outbound

This paper cites MiniCPM4: Ultra-Efficient LLMs on End Devices.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention MiniCPM4: Ultra-Efficient LLMs on End Devices

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.147989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.147989Z digest=sha256:e03a75811f679d0d17d256464399ad465b56665fffcb5f2f3dd64ab396057feb

Observation 99ab8859-6de7-4715-a0e9-96170a4611a9 · outbound

This paper cites Efficient Streaming Language Models with Attention Sinks.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Efficient Streaming Language Models with Attention Sinks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.270498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.270498Z digest=sha256:8504118ee084611e89d076237ced8d95f3eda4f0cde08760a3d1fedc89cd6d86

Observation 73d38bff-7fa9-4a55-90d3-ac941d49c72f · outbound

This paper cites arXiv preprint arXiv:2509.24745 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2509.24745 , year=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.370274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.370274Z digest=sha256:104df771ba32275825192003e36aea60c32f9839ae8535ffda61b8a7e4127877

Observation fd1baf07-ee63-47b5-b399-49b1fe8ac874 · outbound

This paper cites arXiv preprint arXiv:2603.12201 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2603.12201 , year=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.490246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.490246Z digest=sha256:646882b89fd6019a5588d9ddbd7a81258b7e569200cf0561aeee9c61051dd764

Observation e4a60a6e-cf4d-4eaf-80a6-222e9e87a9d3 · outbound

This paper cites arXiv preprint arXiv:2502.18137 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2502.18137 , year=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.590453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.590453Z digest=sha256:ba47f455c3e944133162b401caf353b9536b5e09a2f8559ca5fe35a539ba15e8

Observation e9a3b653-9335-425d-957c-1d984892b036 · outbound

This paper cites arXiv preprint arXiv:2602.13515 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2602.13515 , year=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.706856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.706856Z digest=sha256:ffc69c0021feb904b43c320f6e9fa1232432ed43ea24ff8285c3f41013ce3d67

Observation b61781ed-5b7c-4a77-b6e4-52b028151053 · outbound

This paper cites arXiv preprint arXiv:2603.06199 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2603.06199 , year=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:49.899847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:49.899847Z digest=sha256:6fbd399dbe233e75eb94e06b1948cc06921d9b19e617da4d31876911b7c5c864

Observation 74a91bff-0875-489f-963a-b9c76365960f · outbound

This paper cites Prism: Spectral-Aware Block-Sparse Attention.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Prism: Spectral-Aware Block-Sparse Attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.041954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.041954Z digest=sha256:e5363676af0ddc3d8fc838ccd895a1b468a09b08e887f1dc2d05dd57cf71567e

Observation 78bbb400-e10c-45ef-8c68-c810580fa1f9 · outbound

This paper cites BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention BLASST: Dynamic BLocked Attention Sparsity via Softmax Thresholding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.187273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.187273Z digest=sha256:08522b94550f7028cef6469a0dc4cc7efa280e819a12f87a22a8a613a22c5840

Observation 265fb069-106d-4305-b54a-e46fcc019bc3 · outbound

This paper cites Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Switch Attention: Towards Dynamic and Fine-grained Hybrid Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.320256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.320256Z digest=sha256:620866ef00b1a1509efa646e5fd2df2330074c915fcb84352fc0cb3c88854587

Observation d5b68bfa-d093-4172-9077-d18e87efc085 · outbound

This paper cites arXiv preprint arXiv:2601.17367 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2601.17367 , year=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.398367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.398367Z digest=sha256:1a0c147f167e2e3a4f7f1e25a296f2d630c40eafc2a9ef01f3de9ba9b5eb0849

Observation e75fa642-5e95-44f2-b899-179fb388485d · outbound

This paper cites FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention FlashInfer: Efficient and Customizable Attention Engine for LLM Inference Serving

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.552427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.552427Z digest=sha256:9ef132ebd3f7491eea116fa97bdf1d0e411fdbffa9f653310835745f0810fd73

Observation 4acb7e70-5685-438c-9218-cc5653c26dfb · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Proceedings of the 29th symposium on operating systems principles , pages=

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.676369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.676369Z digest=sha256:17c5f58d6ecae228a7e5a9d79b76daf6f3d9e64fa1b3c43b41fc82a495d52779

Observation 8fca7d37-b64f-4656-bd7d-92f312959700 · outbound

This paper cites Ring Attention with Blockwise Transformers for Near-Infinite Context.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Ring Attention with Blockwise Transformers for Near-Infinite Context

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:50.847935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:50.847935Z digest=sha256:f50b22d18f1a18bd8fb837c4263d9a6bb094630ca91bd0e51b0550e18de5d32a

Observation 4468db8b-d732-4c53-86eb-a825be7d09c6 · outbound

This paper cites International Conference on Learning Representations , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention International Conference on Learning Representations , volume=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.013937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.013937Z digest=sha256:8591a56661ee23463c1493a7353eaa3e91f7291ebe5f0d1584bf0fdd34b1d598

Observation cdbe61bc-a5f0-4be6-8298-18332db0677a · outbound

This paper cites arXiv preprint arXiv:2411.10958 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2411.10958 , year=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.235367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.235367Z digest=sha256:e9d1bb056643ede23e96649c87b871249e41e74180c19d5ce0d22608598f7165

Observation be1ed503-9d27-423b-9bbb-e1333006e8f4 · outbound

This paper cites SageAttention2++: A More Efficient Implementation of SageAttention2.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention SageAttention2++: A More Efficient Implementation of SageAttention2

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.477723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.477723Z digest=sha256:d230fb0a1a6291ab37b1bccabae4f39ce0104fcfef7b20c81a58c3a5da5ec2cd

Observation ffb1585c-97b6-43c4-bfd6-39c2f8acccc9 · outbound

This paper cites 2026 , publisher =.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention 2026 , publisher =

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.646237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.646237Z digest=sha256:594326e70a13f5a19d284b145bc1b2cf692a325428029c0f1c7e3e7b02bf9f72

Observation b3442307-3ab0-4afc-bbf8-9d8b67c8700c · outbound

This paper cites 2023 , publisher =.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention 2023 , publisher =

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.712438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.712438Z digest=sha256:18ecbd71373b95679da21bea98e5648736456f91cf26537b5dbe2e4b48929f7f

Observation 6eb3b274-3a78-486a-9141-8b5d1618cbc3 · outbound

This paper cites GitHub repository , howpublished =.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention GitHub repository , howpublished =

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:51.870523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:51.870523Z digest=sha256:b15b2d0c912b6f6405f7e5569b08d956023ff33b641f4c42cad07e6d0a896c54

Observation 2e7f9c91-1b42-4dc3-b863-4cb021ae7c50 · outbound

This paper cites MiniMax Sparse Attention.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention MiniMax Sparse Attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.016198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.016198Z digest=sha256:b410d27954644099d2878832530a665869edf59c050f9836c2f41601b5e4403b

Observation 372fa9b9-9d05-4a28-becd-5a7d920c12cb · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Advances in Neural Information Processing Systems , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.197288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.197288Z digest=sha256:a4f8097aca2ef29c97c275736c498763925d21fe9dd6d9d5cf5e07d0b19b3378

Observation 7f7801ea-a7bd-4187-89f4-04eb31c3ceff · outbound

This paper cites arXiv preprint arXiv:2602.03560 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2602.03560 , year=

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.348380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.348380Z digest=sha256:7d7782e0743d5846a7396e589a6670215a76470171a539bacb86a8b580323b86

Observation 5cd8f540-4b9a-4f76-84f8-93aaf5fdf7b4 · outbound

This paper cites DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.539170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.539170Z digest=sha256:e764ffecee506f3bd72419f9be3f35849bfdeec5b75bfb7eb6c6f69326eee41c

Observation 752c0eca-ddec-47e4-81c8-83ba42306556 · outbound

This paper cites YaRN: Efficient Context Window Extension of Large Language Models.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention YaRN: Efficient Context Window Extension of Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.698377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.698377Z digest=sha256:e1bfb029339508a11a92052e8989e64ab58019fa1ff54687e6f90ffc44c3a3bd

Observation 9fd9f65d-f319-43ae-8853-6b4db10cc38a · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Kimi K2: Open Agentic Intelligence

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:52.854697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:52.854697Z digest=sha256:431e575f834393a375380c3e50cfeab45ea85723bb89a5949c8ab88e81b38547

Observation cb465c79-47de-4847-bc1e-37be20d626e7 · outbound

This paper cites GLM-5: from Vibe Coding to Agentic Engineering.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention GLM-5: from Vibe Coding to Agentic Engineering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.095194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.095194Z digest=sha256:e1e8be5521e2cea1028bdc041d99836a8512cc1bf04a5201fbe941beb9300fbc

Observation 15ba33b5-eede-4dac-af4a-0abb851a83f2 · outbound

This paper cites Science , volume=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Science , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.255448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.255448Z digest=sha256:8619ba8a4b803b1539fbf8a3b7fab8a36c41ef5643600ff65ad0fff24be8aa6b

Observation c20d5658-82a5-4014-91b5-248a7ddd67fb · outbound

This paper cites Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Retrieval Augmented Generation and Understanding in Vision: A Survey and New Outlook

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.414494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.414494Z digest=sha256:dc4f9f27cd8101ccf6e14809be057a6d95c0a4fb6ca9059bde8115c0656ab365

Observation 962716ca-ffaa-4b15-8dd7-7aff5763e726 · outbound

This paper cites arXiv preprint arXiv:2602.03442 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2602.03442 , year=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.584885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.584885Z digest=sha256:5ec7aae59e8cef1d66aff1f8d87d8c43e9e80602166d0e9a4b10264122672525

Observation 605440b1-68a9-4521-8455-8ac083d751f3 · outbound

This paper cites Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Quest: Query-Aware Sparsity for Efficient Long-Context LLM Inference

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.816311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.816311Z digest=sha256:7e28e8389506e156429b8e7d0fb478f8ca685a6986955734cc93776ef4dd2c70

Observation 5efec613-b623-495e-b49b-da72f5cfa1b5 · outbound

This paper cites Online normalizer calculation for softmax.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Online normalizer calculation for softmax

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:53.999965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:53.999965Z digest=sha256:dfb4a750f9c22019097ecabdc16ebd573952cc6f5c1b5c611128b7e2dbb764bb

Observation 973f1528-763c-4d49-bc90-a226a834dc59 · outbound

This paper cites arXiv preprint arXiv:2510.07486 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2510.07486 , year=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:54.211435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:54.211435Z digest=sha256:925c95f7a1e418e8c39092d6931e65aa97a77a03968bb89902b22898040dfbec

Observation 112c5b92-250a-49a9-9cd3-6567656a2aaa · outbound

This paper cites Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:54.325376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:54.325376Z digest=sha256:bcd7d09a439451d3fc838831bae97f302bd9deea1e10df54515118d4de221334

Observation adac0b8f-8f53-426d-b728-56e26d944b93 · outbound

This paper cites Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:54.484860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:54.484860Z digest=sha256:4a85f14b331fd9c6d55e8fb71c5fe4d624766401bfea65f5e01cb20791807315

Observation 3173efb7-91b2-4634-8514-d4b53d543471 · outbound

This paper cites arXiv preprint arXiv:2603.06274 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2603.06274 , year=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:54.701513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:54.701513Z digest=sha256:8cb950b7dc3db9681eddf01291dd77456f5026c1ae0f4afa31aea7e79df7f7a6

Observation 45f798d9-5313-4f2b-a997-698a771b918a · outbound

This paper cites TokenButler: Token Importance is Predictable.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention TokenButler: Token Importance is Predictable

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:54.858043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:54.858043Z digest=sha256:2d8f2ba17e8816e0cea2c8a9d2cb43dcaed74f32f59de793d9e220939bf10631

Observation d6e48e26-ae9e-49f5-bd9a-1e29d1ebc24e · outbound

This paper cites Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:55.008440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:55.008440Z digest=sha256:9cbcbb453f6af60ba45ea51d99fc6002c8d98bf6f5801d1ee07fc118dc42ab59

Observation 8ad8611a-7984-4426-b3e8-ef3e1b5ac2bc · outbound

This paper cites SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention SSA: Sparse Sparse Attention by Aligning Full and Sparse Attention Outputs in Feature Space

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:55.168098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:55.168098Z digest=sha256:0849872fe2d4e8ddd64059404eff001b4237dc92c78c52f8a6d20b9489c6a66e

Observation 52f282e2-3c08-4793-80a1-46324f01f57c · outbound

This paper cites arXiv preprint arXiv:2512.14082 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2512.14082 , year=

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:55.347696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:55.347696Z digest=sha256:ad9d04c0e39a006ae446b20d5106b31126fa28a1e6b85e9549cada38ec38c3d9

Observation be8b4119-15ce-4dd8-8125-c52d175ddfb2 · outbound

This paper cites arXiv preprint arXiv:2602.05853 , year=.

CoSA: Accelerating Long-Context Inference via Proxy-Kernel Co-Designed Sparse Attention arXiv preprint arXiv:2602.05853 , year=

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-01T02:55:55.511467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T02:55:55.511467Z digest=sha256:859358b0781190f23bd6cfb982b92654c2c0cacdeaba6b1abfdee4e86226b746

Pith citing papers

No inbound Pith citation observations are available.