Pith. sign in

Paper Citation Record · LEDGER

Mask-Based Priors Are More Persistent than Query-Key Initializations

As of 22 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.00418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00418 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:27:28.193940Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fdfd49a-05e7-42ae-9e30-8552cdf81269 · outbound

This paper cites It remains to prove the gradient inequality.

Mask-Based Priors Are More Persistent than Query-Key Initializations It remains to prove the gradient inequality

Reference 1

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:27:28.759504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.189935Z digest=sha256:aaa755675fe7495938c2a705c15b225e17bc1cf3294200a01432b4daef2cf22c

Observation 8764e350-2017-4aa2-8bb5-4297b32f3f38 · outbound

This paper cites Structured initialization for vision transformers.arXiv preprint arXiv:2505.19985,.

Mask-Based Priors Are More Persistent than Query-Key Initializations Structured initialization for vision transformers.arXiv preprint arXiv:2505.19985,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:27:28.607893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.125676Z digest=sha256:5146c65643b635a9e3223f5fcaac62dd4b6a4aada6c07c1d3f36a08e3157e813

Observation 84dc7d6b-e2af-4332-b611-ca63581213db · outbound

This paper cites an unresolved cited work.

Mask-Based Priors Are More Persistent than Query-Key Initializations Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:27:28.746046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.193940Z digest=sha256:056fac8d52e06e0373dac755b0a817d4682c9d66bc083977b1ea2dffed52653a

Observation 5c904b94-a1bc-47d6-b81d-a5fb689638c7 · outbound

This paper cites Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.143731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.143731Z digest=sha256:26bf5ef687682bc8e31ee70c3012a27b87e71d833145ebd0f631b7786d621657

Observation f3805a3d-df84-4147-95be-c71a45f001a4 · outbound

This paper cites Rethinking Causal Mask Attention for Vision-Language Inference.

Mask-Based Priors Are More Persistent than Query-Key Initializations Rethinking Causal Mask Attention for Vision-Language Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.153008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.153008Z digest=sha256:07f6b06c0640de9b359e989a542b1b63839a02ff17981473b7f02f3831158544

Observation 7d14d403-894d-4239-81f4-e7ba3f621644 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.161667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.161667Z digest=sha256:8a512e7bae056a269c7bf76dee441a2ae8e88cce1fe06da31f3e5f776c8f9663

Observation 083545f2-45c0-493a-866a-735402ae473a · outbound

This paper cites k-means mask transformer.

Mask-Based Priors Are More Persistent than Query-Key Initializations k-means mask transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:27:28.773440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.165712Z digest=sha256:fae40e7f59a1e7a0875274e65e4e45ff71c91dfc502ef325cce2f358e0a7290a

Observation 8e401339-a09f-46be-a0de-bec96e36cdf7 · outbound

This paper cites Axial Attention in Multidimensional Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Axial Attention in Multidimensional Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.169464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.169464Z digest=sha256:445a4400b734d55006367f092072050d4d5e1ea3e8b62696220e005907b9caed

Observation 2b3b554b-3fee-4d83-8420-8c817257f1ca · outbound

This paper cites Alex Krizhevsky.

Mask-Based Priors Are More Persistent than Query-Key Initializations Alex Krizhevsky

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.173496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.173496Z digest=sha256:c3dbf459f21f1d7699adf6ea6cd4ea4893f05ec479a4dcb91f8913b3e589bfa6

Observation 1f87a2cb-0bf9-43d0-b72e-1e0f1b53ee10 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Mask-Based Priors Are More Persistent than Query-Key Initializations TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.177064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.177064Z digest=sha256:e2316a8161e4fa99e18cb8fe1190de77b424c8ff3ac0d9b00622246d46c7c12a

Observation ab8a79a0-09b0-4d8c-a339-2f60007d184c · outbound

This paper cites Pointer Sentinel Mixture Models.

Mask-Based Priors Are More Persistent than Query-Key Initializations Pointer Sentinel Mixture Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.180997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.180997Z digest=sha256:4b013dc0e6eba83cacd35713513a2605c50d883ff206aefc65664702f3f52fe4

Observation e1728329-a14a-491e-8ff9-a6be9b4b5aa0 · outbound

This paper cites an unresolved cited work.

Mask-Based Priors Are More Persistent than Query-Key Initializations Unresolved cited work

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.185033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.185033Z digest=sha256:206ec5a60ba47a737068e3c65f7c02100814efdce8fb88f23158a2031fe80076

Observation fd16baa8-a3a4-420d-a838-68a183bb4e3d · outbound

This paper cites Initializing Models with Larger Ones.

Mask-Based Priors Are More Persistent than Query-Key Initializations Initializing Models with Larger Ones

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.129994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.129994Z digest=sha256:add54e0fee5b69dede922a03d68ca5b1b51de0f41e3d2fc51736c8f247809c58

Observation 7e249704-d66a-48a9-b6f7-47cd89494bf5 · outbound

This paper cites Cutting the skip: Training residual-free transformers.arXiv preprint arXiv:2510.00345,.

Mask-Based Priors Are More Persistent than Query-Key Initializations Cutting the skip: Training residual-free transformers.arXiv preprint arXiv:2510.00345,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:27:28.694314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.121006Z digest=sha256:a4d4b81e998459f8bf82193cbc87d618bb4cd47d156652eedbcb94b789995f09

Observation 14d97f61-fe29-40b1-adfe-547b64da66b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mask-Based Priors Are More Persistent than Query-Key Initializations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.106595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.106595Z digest=sha256:815081a74619df94ad51712a2b7908a869302acdf91ed9b188c7344a78e42abd

Observation 70f27395-7d6c-4162-8eb9-d1e0b226714a · outbound

This paper cites Longformer: The Long-Document Transformer.

Mask-Based Priors Are More Persistent than Query-Key Initializations Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.157349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.157349Z digest=sha256:9cd8a45b45fc44316769c5e3f3e46a9ef085616d39f4ab3a68cd8002669fc868

Observation 30c9624f-a04e-407f-900c-307ddc9ec5ef · outbound

This paper cites On weight initialization in deep neural networks.

Mask-Based Priors Are More Persistent than Query-Key Initializations On weight initialization in deep neural networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.116107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.116107Z digest=sha256:20e0b863535a8709658e389170bb28a1757f388760390744130219aaca6a44a9

Observation af3f302e-8d90-4f2a-add0-2eac770f3174 · outbound

This paper cites Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization.

Mask-Based Priors Are More Persistent than Query-Key Initializations Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.134413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.134413Z digest=sha256:ec6449857d71a94537b0e08079957bdff4f40155e33277203e1081cd754e8046

Observation b4230f5a-34a3-4873-8949-d6fc254c08c9 · outbound

This paper cites Weight subcloning: direct initialization of transformers using larger pretrained ones.

Mask-Based Priors Are More Persistent than Query-Key Initializations Weight subcloning: direct initialization of transformers using larger pretrained ones

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.138917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.138917Z digest=sha256:e310601d00583f2b295fb1df314c0f2cfb45f36c88b7b694cdaf6f0563ded271

Observation 8c4228db-04ef-413d-8eaf-0e5e43a801eb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Mask-Based Priors Are More Persistent than Query-Key Initializations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.111257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.111257Z digest=sha256:207b92318b72c3e9ebb243ad1eb8763a89b4b6c8c4fb5014757b9b844218dfd3

Observation d54667f4-4535-4809-a428-b5e0e15eebcf · outbound

This paper cites Jianqiao Zheng, Xueqian Li, and Simon Lucey.

Mask-Based Priors Are More Persistent than Query-Key Initializations Jianqiao Zheng, Xueqian Li, and Simon Lucey

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.149074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.149074Z digest=sha256:c3207dc98eff1e32574ff5f1213ecede24ea7061a1945599848c644ac251a3c4

Pith citing papers

No inbound Pith citation observations are available.