Pith. sign in

Paper Citation Record · LEDGER

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2502.07563.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.07563 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T12:25:42.790098Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.570769Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T18:51:45.610699Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact1
  • verified fuzzy5
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2761e0e-4b64-409e-af90-d7fe4de38bd5 · outbound

This paper cites GPT-4 Technical Report.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.612941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.612941Z digest=sha256:d4e9c29aab2afa03b0a397ed4f3a80a24f9e90321be98eae35b5f91ae9ac503f

Observation d1565af2-adb7-4c74-b980-67bf83a13cfe · outbound

This paper cites Linear Transformers with Learnable Kernel Functions are Better In-Context Models.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Linear Transformers with Learnable Kernel Functions are Better In-Context Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.624361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.624361Z digest=sha256:2a2384d31bd9d4009725d6238c05ddb410660ca92fbeff71aa59bc1339700aa3

Observation 3be2ee39-3959-4271-bd90-f009297dddef · outbound

This paper cites Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-08T12:25:43.257935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.635272Z digest=sha256:5e295fdccb6512226040a206b52c9d187a90a0dd673184c392c106875f195a02

Observation 3ab4cdec-02e3-43ca-8111-9d3670376504 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.640835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.640835Z digest=sha256:c7298776f7609d80931f316dcd1d659bdb66f5eb90c2d83d0eff93983bb07beb

Observation 91312a3a-7ed3-4b88-866a-fea054710de2 · outbound

This paper cites Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.646406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.646406Z digest=sha256:e7ec5f885d7c12d116141b91c3b657c4634f458f20d7681c3e0f0badf82ddeac

Observation c925f4d6-1bc0-4283-880d-3c95e3e0f5c8 · outbound

This paper cites The Llama 3 Herd of Models.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.656382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.656382Z digest=sha256:31262a9dc7e12bdad0612166cdd4994bd7a88691b89e97e88f22162c0660e9bb

Observation b7c14e75-2bbc-48cd-878f-db7fdcfb4804 · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.661139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.661139Z digest=sha256:01202b5609da32a93b3250cea7334a8668b19e44e72c5b0443f09fa05cd3c5f6

Observation 834dc8f1-c38c-4d0e-a4bc-f1dc493a1ecc · outbound

This paper cites Measuring Massive Multitask Language Understanding.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Measuring Massive Multitask Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.666217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.666217Z digest=sha256:baf6c490f59a9c56a22343fdc06a43d4cd6e8dca0d81c858c1adcf48442274af

Observation cc935757-5111-4368-955a-85e43c483e88 · outbound

This paper cites Repeat After Me: Transformers are Better than State Space Models at Copying.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Repeat After Me: Transformers are Better than State Space Models at Copying

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.671185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.671185Z digest=sha256:1c762b22ba806b61f1af315163bba7828a7add6816db0ff995223dbc14d9b8c1

Observation 2ffb14d9-494f-4d45-a10e-7eb99a1d6501 · outbound

This paper cites MiniMax-01: Scaling Foundation Models with Lightning Attention.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid MiniMax-01: Scaling Foundation Models with Lightning Attention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.680742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.680742Z digest=sha256:7d7530140c91120366ac97bd92d98d2d2856f06e608a44ec2733706b4c95e1e6

Observation b69a9752-719a-47e6-bd6d-a2636c0c7e99 · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Jamba: A Hybrid Transformer-Mamba Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.685122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.685122Z digest=sha256:b8d6fcca823a805cbd85ac346033a4020ba92db3d44429d7cdf6c5691d4fed31

Observation cf67f9af-7e53-47d2-bdae-d883706bb468 · outbound

This paper cites an unresolved cited work.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T12:25:43.422421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.689644Z digest=sha256:5257841a1697c7361fdb2d1cc4e5994fa8493c00f7f5dcd46c632e587a6e1a69

Observation 12a6d39f-d016-4439-87ba-0d22cc0d3524 · outbound

This paper cites doi: 10.18653/v1/2023.findings-emnlp.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid doi: 10.18653/v1/2023.findings-emnlp

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.693907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.693907Z digest=sha256:1f5a3fa65044f7b3289a37bd6c319d262b593f174aea285966f001fcb70679e1

Observation 7c2fab43-7087-4322-8cde-93c3f06e6345 · outbound

This paper cites Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.703293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.703293Z digest=sha256:cad41c9b032336a23db0cdc35ad15b5b2c285eb7c4fc6539760189b3d72e61ff

Observation 7b8e93ba-727c-40d1-abd2-01c8cd362928 · outbound

This paper cites TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.707448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.707448Z digest=sha256:125dc17d73a04f404a666ef2ad984450af1d06f314a82e7411506a569a5810ac

Observation 339ae241-9a41-4f7a-9c8e-ecafd7c78a05 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.711865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.711865Z digest=sha256:31cdf5bbe6dcc54046657a080da0ab52cb72913bb594b1d6aec16a9df70d2197

Observation ee0d4a49-1b80-4720-9669-95881963ddff · outbound

This paper cites Scaling Laws for Linear Complexity Language Models.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Scaling Laws for Linear Complexity Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.721108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.721108Z digest=sha256:3ff10deff0253edac0b60b69fb24272cc4894a91dc11a096d13c53a8cfe8d661

Observation 88ccc2bf-1a61-4aed-8e0d-f1c7062771e4 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.725885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.725885Z digest=sha256:f44782a207d61ae024ac402352818fe57021c3a2c46c99018e881f12451b3168

Observation 7ec65b29-6216-4b59-9f1b-0d3a79743ac3 · outbound

This paper cites Linear Attention Sequence Parallelism.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Linear Attention Sequence Parallelism

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T12:25:42.952657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.730995Z digest=sha256:b2137b5c81c9b0bbd9b4c9b70168d6c15b17d0403bf09d3e2e6e2efc98ceb694

Observation ab372280-d145-4bd1-91db-11abde4aaaf2 · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.741468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.741468Z digest=sha256:c9169895d0dbb818874a13c359bd8dde3ea30446e663976916345b09898507d9

Observation c0ec0ec4-683c-44c3-ae7b-462d4e62bebd · outbound

This paper cites Gated Linear Attention Transformers with Hardware-Efficient Training.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Gated Linear Attention Transformers with Hardware-Efficient Training

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.746310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.746310Z digest=sha256:defec111fb9f735c1a1a344fac57e3449b41d91c7ef719e0cbd15679c108ed2b

Observation a2b0d8cd-b9c5-450a-8bba-bde9b09a046c · outbound

This paper cites Parallelizing Linear Transformers with the Delta Rule over Sequence Length.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Parallelizing Linear Transformers with the Delta Rule over Sequence Length

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.751276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.751276Z digest=sha256:9bdb8f930a1bd22ceddfa60fdbf708e9a40736997a172ffa6a291f1e2622abab

Observation b7a65e53-4525-4a5c-b54c-566a1621e8df · outbound

This paper cites Boosting Distributed Training Performance of the Unpadded BERT Model.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Boosting Distributed Training Performance of the Unpadded BERT Model

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.756424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.756424Z digest=sha256:3b2fb4a3883d56206bd60e9cf4d6951c50975de4c99a9eed2352602fac77fb90

Observation 6760dd63-efe1-42db-ba9a-47bb5f14016b · outbound

This paper cites ByteTransformer: A high- performance transformer boosted for variable-length in- puts.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid ByteTransformer: A high- performance transformer boosted for variable-length in- puts

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:25:43.406108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.761073Z digest=sha256:da62870fd11b801ace36487958dff58df1e097350d596fa3980d676a18049d9c

Observation 63a8ebcd-779d-4711-b25d-16f3495823ca · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.771337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.771337Z digest=sha256:3279a462d3462d5386d8694463b3ef35ea8d6116e25f0e120cc7edd691f314cd

Observation 69562057-6f52-486b-a7b7-5ec38ad4514f · outbound

This paper cites As these techniques are variants of data parallelism, they integrate seamlessly with LASP.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid As these techniques are variants of data parallelism, they integrate seamlessly with LASP

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:25:43.389640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.776156Z digest=sha256:5e2c8333944853794b5cc123866e73b6940dd5600084e0ef895add58c0783064

Observation 1571c3da-451f-4f1c-9c49-dcd8957877ae · outbound

This paper cites LASP-2 can manage variable sequence lengths efficiently by treating the entire batch as a single long sequence, streamlining the process without requiring padding.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LASP-2 can manage variable sequence lengths efficiently by treating the entire batch as a single long sequence, streamlining the process without requiring padding

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:25:43.373969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.781266Z digest=sha256:76acf16682ee46d257f8c626d245aa4bcff2f124e93f78ca25d0cfb0545338c2

Observation 3e494f39-e4b1-477b-881e-934873950d53 · outbound

This paper cites Split Size of Gathering 2048 512 128 32 Number of Splits 1 4 16 64 Throughput 486183 486166 486169 486158 A.5.4.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Split Size of Gathering 2048 512 128 32 Number of Splits 1 4 16 64 Throughput 486183 486166 486169 486158 A.5.4

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:25:43.340423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.790098Z digest=sha256:c5eead6708bda2285f6d522bcc2f327abb40b6467727f7314a741711517b864a

Observation d3b1d05e-81bf-411b-bf0b-c8b773864e59 · outbound

This paper cites Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence

Reference 936

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.698693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.698693Z digest=sha256:973e4c8c87b93b76992e163e7a789d4243e43437d9bf1456bdacd1aaba552310

Observation c48c134d-9f26-4c62-86c9-29f0f81fa76a · outbound

This paper cites An Empirical Study of Mamba-based Language Models.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid An Empirical Study of Mamba-based Language Models

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.735960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.735960Z digest=sha256:a6d3f7d4351404e7df00fa91091fa179980b51e82335ad2d29642936ec317a7d

Observation 32bc18e4-d811-4504-82ba-4ea714b00317 · outbound

This paper cites Gated Slot Attention for Efficient Linear-Time Sequence Modeling.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Gated Slot Attention for Efficient Linear-Time Sequence Modeling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.766684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.766684Z digest=sha256:445f6d5e886b93a96ff561e94a23b4260ae77b6040f552020590a95d9a59712f

Observation a2c982bb-f607-469c-9c02-db98ad7db2a2 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Adam: A Method for Stochastic Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.676350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.676350Z digest=sha256:54f34cef8b6f55c32219543e3e113020fda54d4813fab1c6807bf0439c06283e

Observation 98267f49-1450-4e61-b3ab-b1c491650e7d · outbound

This paper cites FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.716501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.716501Z digest=sha256:59c2ed54cf0b346c2b66022be615d7f32cb5dc5235d55537ff39306aee27722e

Observation 685225d0-e7d6-4c7b-a73b-8e6e7070c9d1 · outbound

This paper cites Fewer Truncations Improve Language Modeling.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Fewer Truncations Improve Language Modeling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.651269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.651269Z digest=sha256:cc5ad48be3faad621dcf677fec177d6236b0c28704d524e0dab5a5bc80040bf8

Observation c6aee1e7-a230-4000-80da-79992ce586cb · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.618676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.618676Z digest=sha256:6822a4e013ceea7f5d41ced20585baa7b67a38ed56a256777e1d06b8e8470314

Observation 43ea1b62-b1b9-43a9-8ba9-e0a1cbe80823 · outbound

This paper cites Simple linear attention language models balance the recall-throughput tradeoff.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Simple linear attention language models balance the recall-throughput tradeoff

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T12:25:42.629773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:25:42.629773Z digest=sha256:b1340bc48145208732758861e32331cdf78ee3078b33f81da9657262c72f555a

Observation ca173ad3-351c-43fc-9841-3ce85014a41d · outbound

This paper cites L" denotes linear Transformer layers and.

LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid L" denotes linear Transformer layers and

Reference 2048

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T12:25:43.357187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T12:25:42.785686Z digest=sha256:068e9ba586595382e874bb3b7165dd06ed6df5073ca88da99478de0e20789d4f

Pith citing papers

Observation 78e836d7-64ee-4e55-ab47-b3727d46e310 · inbound

SpikingBrain: Spiking Brain-inspired Large Models cites this paper.

SpikingBrain: Spiking Brain-inspired Large Models LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T18:51:45.613221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T18:51:06.243305Z digest=sha256:475399ec0bbce630b8d1b4e3c6132c3810988f459d5eb488efe4ebcd345d825b

Observation 2967adc6-a89e-4b02-aa21-be8e8d1c471d · inbound

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix cites this paper.

A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T17:32:37.570769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T17:32:37.570769Z digest=sha256:e6352ee46e7ba871d208288afd40d1764d366ba0a913eb1a954dc6178cc915ad