Pith. sign in
Pith Number

pith:5BWLUGJF

pith:2026:5BWLUGJFIWJUYIIVLYW3DPPQJP
not attested not anchored not stored refs resolved

Full Attention Strikes Back: Transferring Full Attention into Sparse within Hundred Training Steps

Hanlin Tang, Kan Liu, Lan Tao, Lin Qu, Maohua Li, Xiaoxing Ma, Yanke Zhou, Yiduo Li, Yuan Yao

Full-attention LLMs already contain the structure to become highly sparse models after only a few hundred training steps.

arxiv:2605.16928 v1 · 2026-05-16 · cs.CL · cs.AI

Add to your LaTeX paper
\usepackage{pith}
\pithnumber{5BWLUGJFIWJUYIIVLYW3DPPQJP}

Prints a linked badge after your title and injects PDF metadata. Compiles on arXiv. Learn more · Embed verified badge

Record completeness

1 Bitcoin timestamp
2 Internet Archive
3 Author claim open · sign in to claim
4 Citations open
5 Replications open
Portable graph bundle live · download bundle · merged state
The bundle contains the canonical record plus signed events. A mirror can host it anywhere and recompute the same current state with the deterministic merge algorithm.

Claims

C1strongest claim

full-attention LLMs are already intrinsically sparse and can be transformed into highly sparse models with only minimal adaptation

C2weakest assumption

only a small subset of attention heads truly requires full long-context processing and long-range retrieval is governed primarily by a low-dimensional subspace

C3one line summary

RTPurbo exploits intrinsic sparsity in full-attention LLMs to achieve near-lossless sparse inference after only a few hundred training steps via retrieval-head identification and a lightweight token indexer.

References

34 extracted · 34 resolved · 9 Pith anchors

[1] In: Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (Aug 2024).https://doi.org/10.18653/v1/2024.acl-long.172 2024 · doi:10.18653/v1/2024.acl-long.172
[2] Matharena: Evaluating llms on uncontaminated math competitions, February 2025 2025
[3] Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities 2025 · arXiv:2507.06261
[4] Flashattention-2: Faster attention with better parallelism and work partitioning 2024
[5] DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models 2025 · arXiv:2512.02556

Formal links

2 machine-checked theorem links

Cited by

1 paper in Pith

Receipt and verification
First computed 2026-05-20T00:03:31.256999Z
Builder pith-number-builder-2026-05-17-v1
Signature Pith Ed25519 (pith-v1-2026-05) · public key
Schema pith-number/v1.0

Canonical hash

e86cba192545934c21155e2db1bdf04bce33534c9fdd005f3e77d76b9e75406e

Aliases

arxiv: 2605.16928 · arxiv_version: 2605.16928v1 · doi: 10.48550/arxiv.2605.16928 · pith_short_12: 5BWLUGJFIWJU · pith_short_16: 5BWLUGJFIWJUYIIV · pith_short_8: 5BWLUGJF
Agent API
Verify this Pith Number yourself
curl -sH 'Accept: application/ld+json' https://pith.science/pith/5BWLUGJFIWJUYIIVLYW3DPPQJP \
  | jq -c '.canonical_record' \
  | python3 -c "import sys,json,hashlib; b=json.dumps(json.loads(sys.stdin.read()), sort_keys=True, separators=(',',':'), ensure_ascii=False).encode(); print(hashlib.sha256(b).hexdigest())"
# expect: e86cba192545934c21155e2db1bdf04bce33534c9fdd005f3e77d76b9e75406e
Canonical record JSON
{
  "metadata": {
    "abstract_canon_sha256": "5503d7ff56dea961efbb7aa31cfe7fa14d7c260239051500288a33744cc2cb34",
    "cross_cats_sorted": [
      "cs.AI"
    ],
    "license": "http://creativecommons.org/licenses/by/4.0/",
    "primary_cat": "cs.CL",
    "submitted_at": "2026-05-16T10:51:58Z",
    "title_canon_sha256": "eb90e3c553c5f824f187993fb0f75f6152125c643a405384461d90826ccf1567"
  },
  "schema_version": "1.0",
  "source": {
    "id": "2605.16928",
    "kind": "arxiv",
    "version": 1
  }
}