Pith. sign in

Paper Citation Record · LEDGER

Mask-Based Priors Are More Persistent than Query-Key Initializations

As of 21 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2608.00418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00418 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:27:28.193940Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact2
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fdfd49a-05e7-42ae-9e30-8552cdf81269 · outbound

This paper cites It remains to prove the gradient inequality.

Mask-Based Priors Are More Persistent than Query-Key Initializations It remains to prove the gradient inequality

Reference 1

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:27:28.759504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.189935Z digest=sha256:0eb75c4a88a0d87496dd54ca7a1122fba7e484c8b340ed6196f52bbce1a03572

Observation 8764e350-2017-4aa2-8bb5-4297b32f3f38 · outbound

This paper cites Structured initialization for vision transformers.arXiv preprint arXiv:2505.19985,.

Mask-Based Priors Are More Persistent than Query-Key Initializations Structured initialization for vision transformers.arXiv preprint arXiv:2505.19985,

Reference 5

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:27:28.607893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.125676Z digest=sha256:d8878a2531d9f0e0236a02c02ffde42e146347ac64e0e2a74a3f41f6c20f7649

Observation 84dc7d6b-e2af-4332-b611-ca63581213db · outbound

This paper cites an unresolved cited work.

Mask-Based Priors Are More Persistent than Query-Key Initializations Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:27:28.746046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.193940Z digest=sha256:a8223a6e3024eb239642a1bc5e6c61b207f074f7580ca078c699bb3469eb6048

Observation 5c904b94-a1bc-47d6-b81d-a5fb689638c7 · outbound

This paper cites Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Algorithmic Task Capture, Computational Complexity, and Inductive Bias of Infinite Transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.143731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.143731Z digest=sha256:1a39e6f134c2d395f80ddd933fc1ec3ce1c19283bf85bcf92b59a77f8d4aad6c

Observation f3805a3d-df84-4147-95be-c71a45f001a4 · outbound

This paper cites Rethinking Causal Mask Attention for Vision-Language Inference.

Mask-Based Priors Are More Persistent than Query-Key Initializations Rethinking Causal Mask Attention for Vision-Language Inference

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.153008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.153008Z digest=sha256:8f562865e2b294898981220528192fed4945c48ed4aeb8656350418871d1843c

Observation 7d14d403-894d-4239-81f4-e7ba3f621644 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Generating Long Sequences with Sparse Transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.161667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.161667Z digest=sha256:2327a1a04ba4ffafe2a8f17618e397fd0488b1a9de3f7dd0be801ff821dadc4d

Observation 083545f2-45c0-493a-866a-735402ae473a · outbound

This paper cites k-means mask transformer.

Mask-Based Priors Are More Persistent than Query-Key Initializations k-means mask transformer

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:27:28.773440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.165712Z digest=sha256:37adaf6f5a2601f0f7553502bbdfebc7c36ebf33303265e325e6315d7a4c0e4e

Observation 8e401339-a09f-46be-a0de-bec96e36cdf7 · outbound

This paper cites Axial Attention in Multidimensional Transformers.

Mask-Based Priors Are More Persistent than Query-Key Initializations Axial Attention in Multidimensional Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.169464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.169464Z digest=sha256:8df3d87721a4b83858089d83b03951f9533743a6c97f1963e8535196aa9949ce

Observation 2b3b554b-3fee-4d83-8420-8c817257f1ca · outbound

This paper cites Alex Krizhevsky.

Mask-Based Priors Are More Persistent than Query-Key Initializations Alex Krizhevsky

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.173496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.173496Z digest=sha256:020a121ce924d558f08eb09ed6b19ee6768375bc487a54674ac46526786e0d9e

Observation 1f87a2cb-0bf9-43d0-b72e-1e0f1b53ee10 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

Mask-Based Priors Are More Persistent than Query-Key Initializations TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.177064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.177064Z digest=sha256:35c2611b4088a16ec95e46741c52b33bf6a1873aa8b9b8c2eb7baac80f83cc69

Observation ab8a79a0-09b0-4d8c-a339-2f60007d184c · outbound

This paper cites Pointer Sentinel Mixture Models.

Mask-Based Priors Are More Persistent than Query-Key Initializations Pointer Sentinel Mixture Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.180997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.180997Z digest=sha256:5fac23bc527e5276b7a74431da20cc4206e2a87306ccdbd85ca7f3201ae5e360

Observation e1728329-a14a-491e-8ff9-a6be9b4b5aa0 · outbound

This paper cites an unresolved cited work.

Mask-Based Priors Are More Persistent than Query-Key Initializations Unresolved cited work

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.185033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.185033Z digest=sha256:977e47f9a6109369b7d472ed64c6a1c39727a350640de38f83f6d782ba147faa

Observation fd16baa8-a3a4-420d-a838-68a183bb4e3d · outbound

This paper cites Initializing Models with Larger Ones.

Mask-Based Priors Are More Persistent than Query-Key Initializations Initializing Models with Larger Ones

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.129994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.129994Z digest=sha256:ebb56127b901da4a7b79d86089d4da351618e479b863906c3ded40044acac39d

Observation 7e249704-d66a-48a9-b6f7-47cd89494bf5 · outbound

This paper cites Cutting the skip: Training residual-free transformers.arXiv preprint arXiv:2510.00345,.

Mask-Based Priors Are More Persistent than Query-Key Initializations Cutting the skip: Training residual-free transformers.arXiv preprint arXiv:2510.00345,

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:27:28.694314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-15T15:27:28.121006Z digest=sha256:083512a6a651bd3b655224d59ea24ef35a277cdf743fac0032f9c7210e90a3a5

Observation 14d97f61-fe29-40b1-adfe-547b64da66b1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Mask-Based Priors Are More Persistent than Query-Key Initializations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.106595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.106595Z digest=sha256:585c7b8f81dc2235447ea4be6dd9ae9a84ce8f326628ce0648c19ca74f10a615

Observation 70f27395-7d6c-4162-8eb9-d1e0b226714a · outbound

This paper cites Longformer: The Long-Document Transformer.

Mask-Based Priors Are More Persistent than Query-Key Initializations Longformer: The Long-Document Transformer

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.157349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.157349Z digest=sha256:07dfa0b95ee252d75802d0ce0e31738bd6cd448870b321359a64d70463bfe97a

Observation 30c9624f-a04e-407f-900c-307ddc9ec5ef · outbound

This paper cites On weight initialization in deep neural networks.

Mask-Based Priors Are More Persistent than Query-Key Initializations On weight initialization in deep neural networks

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.116107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.116107Z digest=sha256:a5c1f019e20d0e561ec8dce431cb6faeccb62da29f85150950e1d1be3fbdcd75

Observation af3f302e-8d90-4f2a-add0-2eac770f3174 · outbound

This paper cites Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization.

Mask-Based Priors Are More Persistent than Query-Key Initializations Scaling Smart: Accelerating Large Language Model Pre-training with Small Model Initialization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.134413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.134413Z digest=sha256:e6c22de8d7b45d56b79867b9b5ebc86f2e5f99814a3cb6f9dabe17c9d83abef0

Observation b4230f5a-34a3-4873-8949-d6fc254c08c9 · outbound

This paper cites Weight subcloning: direct initialization of transformers using larger pretrained ones.

Mask-Based Priors Are More Persistent than Query-Key Initializations Weight subcloning: direct initialization of transformers using larger pretrained ones

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.138917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.138917Z digest=sha256:c398f18171fa3d9eb5b91d511c1a8c9004391678929b29d8f626f981fd132b1d

Observation 8c4228db-04ef-413d-8eaf-0e5e43a801eb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Mask-Based Priors Are More Persistent than Query-Key Initializations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.111257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.111257Z digest=sha256:e7eebd7d00814a847ee8d86706eafa0829647bedc6c6814144e7791ab452fa86

Observation d54667f4-4535-4809-a428-b5e0e15eebcf · outbound

This paper cites Jianqiao Zheng, Xueqian Li, and Simon Lucey.

Mask-Based Priors Are More Persistent than Query-Key Initializations Jianqiao Zheng, Xueqian Li, and Simon Lucey

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-15T15:27:28.149074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:27:28.149074Z digest=sha256:e7d776f987c04b8ae972a8b973d030903b88aa93fed16f7dfb695e3494ba15fa

Pith citing papers

No inbound Pith citation observations are available.