Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:25:42.790098Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2502.07563.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T12:25:42.790098Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T17:32:37.570769Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T18:51:45.610699Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b2761e0e-4b64-409e-af90-d7fe4de38bd5 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1565af2-adb7-4c74-b980-67bf83a13cfe · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Linear Transformers with Learnable Kernel Functions are Better In-Context Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3be2ee39-3959-4271-bd90-f009297dddef · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Searching for Needles in a Haystack: On the Role of Incidental Bilingualism in PaLM's Translation Capability
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ab4cdec-02e3-43ca-8111-9d3670376504 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91312a3a-7ed3-4b88-866a-fea054710de2 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c925f4d6-1bc0-4283-880d-3c95e3e0f5c8 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7c14e75-2bbc-48cd-878f-db7fdcfb4804 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 834dc8f1-c38c-4d0e-a4bc-f1dc493a1ecc · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Measuring Massive Multitask Language Understanding
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc935757-5111-4368-955a-85e43c483e88 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Repeat After Me: Transformers are Better than State Space Models at Copying
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ffb14d9-494f-4d45-a10e-7eb99a1d6501 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid MiniMax-01: Scaling Foundation Models with Lightning Attention
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b69a9752-719a-47e6-bd6d-a2636c0c7e99 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Jamba: A Hybrid Transformer-Mamba Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf67f9af-7e53-47d2-bdae-d883706bb468 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12a6d39f-d016-4439-87ba-0d22cc0d3524 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid doi: 10.18653/v1/2023.findings-emnlp
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c2fab43-7087-4322-8cde-93c3f06e6345 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Dataset Decomposition: Faster LLM Training with Variable Sequence Length Curriculum
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b8e93ba-727c-40d1-abd2-01c8cd362928 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid TransNormerLLM: A Faster and Better Large Language Model with Improved TransNormer
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 339ae241-9a41-4f7a-9c8e-ecafd7c78a05 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0d4a49-1b80-4720-9669-95881963ddff · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Scaling Laws for Linear Complexity Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88ccc2bf-1a61-4aed-8e0d-f1c7062771e4 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ec65b29-6216-4b59-9f1b-0d3a79743ac3 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Linear Attention Sequence Parallelism
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab372280-d145-4bd1-91db-11abde4aaaf2 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LongVILA: Scaling Long-Context Visual Language Models for Long Videos
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0ec0ec4-683c-44c3-ae7b-462d4e62bebd · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Gated Linear Attention Transformers with Hardware-Efficient Training
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2b0d8cd-b9c5-450a-8bba-bde9b09a046c · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Parallelizing Linear Transformers with the Delta Rule over Sequence Length
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a65e53-4525-4a5c-b54c-566a1621e8df · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Boosting Distributed Training Performance of the Unpadded BERT Model
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6760dd63-efe1-42db-ba9a-47bb5f14016b · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid ByteTransformer: A high- performance transformer boosted for variable-length in- puts
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 63a8ebcd-779d-4711-b25d-16f3495823ca · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69562057-6f52-486b-a7b7-5ec38ad4514f · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid As these techniques are variants of data parallelism, they integrate seamlessly with LASP
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1571c3da-451f-4f1c-9c49-dcd8957877ae · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid LASP-2 can manage variable sequence lengths efficiently by treating the entire batch as a single long sequence, streamlining the process without requiring padding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3e494f39-e4b1-477b-881e-934873950d53 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Split Size of Gathering 2048 512 128 32 Number of Splits 1 4 16 64 Throughput 486183 486166 486169 486158 A.5.4
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d3b1d05e-81bf-411b-bf0b-c8b773864e59 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Eagle and Finch: RWKV with Matrix-Valued States and Dynamic Recurrence
Reference 936
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c48c134d-9f26-4c62-86c9-29f0f81fa76a · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid An Empirical Study of Mamba-based Language Models
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32bc18e4-d811-4504-82ba-4ea714b00317 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Gated Slot Attention for Efficient Linear-Time Sequence Modeling
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2c982bb-f607-469c-9c02-db98ad7db2a2 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Adam: A Method for Stochastic Optimization
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98267f49-1450-4e61-b3ab-b1c491650e7d · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-precision
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 685225d0-e7d6-4c7b-a73b-8e6e7070c9d1 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Fewer Truncations Improve Language Modeling
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6aee1e7-a230-4000-80da-79992ce586cb · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43ea1b62-b1b9-43a9-8ba9-e0a1cbe80823 · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid Simple linear attention language models balance the recall-throughput tradeoff
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca173ad3-351c-43fc-9841-3ce85014a41d · outbound
LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid L" denotes linear Transformer layers and
Reference 2048
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 78e836d7-64ee-4e55-ab47-b3727d46e310 · inbound
SpikingBrain: Spiking Brain-inspired Large Models LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2967adc6-a89e-4b02-aa21-be8e8d1c471d · inbound
A Training-Memory Regression in MLA Sequence Parallelism: Why Megatron-Core Forbids Absorption, and LAGA -- a Communication-Efficient Fix LASP-2: Rethinking Sequence Parallelism for Linear Attention and Its Hybrid
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.