Pith. sign in

Paper Citation Record · LEDGER

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

As of 7 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.19531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19531 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:21:02.060803Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:57:10.558525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfeadc09-5c90-4514-a781-34d65330052a · outbound

This paper cites Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.129826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.129826Z digest=sha256:815bcff18b94f75f2614300a0097c0cb49aab43f33b2b41e083895fbc3ac165f

Observation 3613ede7-6236-4cd7-b96b-3bef5ecbc367 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.303448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.303448Z digest=sha256:39a9e745955a8424b03cd7365d16a363c773bf6e21b864a0fc7b1e83b3d0a1fe

Observation 048926cc-308b-46de-867a-170a7ef15673 · outbound

This paper cites Circuit Complexity Bounds for RoPE-based Transformer Architecture.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Circuit Complexity Bounds for RoPE-based Transformer Architecture

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:03.091343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:00.473856Z digest=sha256:d6500f75f079c8135c47ab450e33cebbeb64da33f7a2620c4c2e359088555512

Observation b244969a-2e56-4356-b7ec-223e8ae722a6 · outbound

This paper cites Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.912542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:00.530502Z digest=sha256:3a98301e90304deefbcf2b7886e85d20008004fc4dd8c13d80c4f19c0647e22b

Observation d2393cb8-fdb7-4118-bac0-38c4439db750 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.641603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.641603Z digest=sha256:40693d4a010d09404b33afc2d145e5c8228604b33964641c100785e3d3b74550

Observation 98c36a3c-2de1-4115-adcb-e61fcb878086 · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.809131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.809131Z digest=sha256:32a143a6676a5be18559f8ed05ffe012ab3eb8d0be7cefa3be2c8bd949a37e5c

Observation 063b1293-4e99-4560-8953-4c1d2dd07be4 · outbound

This paper cites Subquadratic Algorithms and Hardness for Attention with Any Temperature.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Subquadratic Algorithms and Hardness for Attention with Any Temperature

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.673064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:00.874738Z digest=sha256:2957f74f88b36b45d3177af819523c89abfaa73d0fe260450ec5096c9a19aa32

Observation 56df5c5a-39b4-4178-b882-02195e1018e8 · outbound

This paper cites In-Context Convergence of Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions In-Context Convergence of Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.978836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.978836Z digest=sha256:654643172cbd5e9251407bcebeafa2f5d64d776ba5d3c0404e1de6d646fbf1d6

Observation f226550d-9ed6-48ea-858b-4d98beffd732 · outbound

This paper cites Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.293029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.293029Z digest=sha256:95895345a437299eb9170ee284746895bbd885d957980c8ef4726e07425c1dc1

Observation e4caa9f6-825f-4547-9a72-b008be475c2d · outbound

This paper cites Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.435527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:01.340668Z digest=sha256:6310ca944e2ae6257562a2b53175949d9d05625b5a03b1b082b2e41ccd695c01

Observation 15190e46-7116-43bb-bada-dbb6eb711a62 · outbound

This paper cites On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.419305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.419305Z digest=sha256:bb1cc987dad02ffef9c206f32968de7ff137609d95d8edcc9d92db80c74911c1

Observation 1027cf28-3136-40e0-9c9b-44a8bdd5b271 · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.467347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.467347Z digest=sha256:bbb51a0bf8b1dc18bfc11b0616c206643593323cf11a9213a5858db3963d3a30

Observation 9cadd4b5-c628-4d76-9b40-56a422be704e · outbound

This paper cites Evalu- ating numeracy of language models as a natural language inference task.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Evalu- ating numeracy of language models as a natural language inference task

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:21:03.534648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:01.519168Z digest=sha256:832968b0fd13b8e013437c3d0540b1be5120191dd44c136b0a1fbf5b8e8c945d

Observation a5275ba0-4a72-438a-b059-5c7f92ff82ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.699323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.699323Z digest=sha256:1d11dcc192f8160d37e8e1867cc6aaf5d9a44e3da041a2ec20282d2ba9396e9b

Observation 595b73b5-7fae-49f6-895d-3b764859bc8a · outbound

This paper cites Transformers Learn Low Sensitivity Functions: Investigations and Implications.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers Learn Low Sensitivity Functions: Investigations and Implications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.743928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.743928Z digest=sha256:fea70f10bd34bc36bbfcc08afcac2755049edbfb8107351b38e31ddbb251728d

Observation e3256b49-7b24-4451-9d9e-d7816e3bed41 · outbound

This paper cites Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.786208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.786208Z digest=sha256:82a2d322aab7e3093beb623003eb96cf034eb00a8e816396a6e76df5d080356b

Observation 939fe7ff-cdb3-4f8f-88fc-e807cdd348f8 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.863974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.863974Z digest=sha256:672ae55d45fe9e7d54a0dfdd71888a7c92dc8037017a546c87413d5f7e91e17c

Observation 700e4e8a-caf4-45ef-bfb2-347f4c249a7d · outbound

This paper cites DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.955998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.955998Z digest=sha256:41a471bc4531eb8d0ce3d59d11ec60d71cb9c249a0a6c3311b504ab871d39a3c

Observation 64b3c3ce-f5d0-49a0-bfd9-f86abd0ff6ae · outbound

This paper cites Genomeocean: An efficient genome foundation model trained on large-scale metagenomic assemblies.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Genomeocean: An efficient genome foundation model trained on large-scale metagenomic assemblies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:21:03.344204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T14:21:01.993859Z digest=sha256:a4e2b0aeaa79e0964e51292af19f177c1ec9b552c4a7ad6694f8951c2377c9aa

Observation fd78cc70-0225-422b-85bf-579607168ac3 · outbound

This paper cites DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:02.060803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:02.060803Z digest=sha256:83317a84e98ab269408ce97e11d247f242ce0ecea08157ef5182f09410cf3685

Observation f0e4350d-26ba-4df3-b90e-ba5b1d39be14 · outbound

This paper cites Transformers in Uniform TC$^0$.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers in Uniform TC$^0$

Reference 1867

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.402417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.402417Z digest=sha256:55a34d32d97ac49dc960176f72c9a502ea81c32ffc218a7db79bfa2c83615f8f

Observation 206e2655-941d-4e3c-a85b-05d3558c2bba · outbound

This paper cites Why are Sensitive Functions Hard for Transformers?.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Why are Sensitive Functions Hard for Transformers?

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.050664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.050664Z digest=sha256:b971e4f642cb6a58f1b99c91bb87013c329571560c05834725d59631c9391717

Observation 23287124-bca5-4cf0-8d95-87a8944e8a85 · outbound

This paper cites Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.742146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.742146Z digest=sha256:935144f8ceab1b1f4682edfd3907bff0c0a277e9a29e170ff822b5be80cf3fe6

Observation 826301cb-e573-4121-bba2-e1db4ad29938 · outbound

This paper cites Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.122910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.122910Z digest=sha256:9f6b12148ab1ea9490b031e78a90abca7fa844ffb8f9da17d338d51328d3b659

Observation ba15c0a1-cce9-4037-83c4-9e7db19ab1dc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions LLaMA: Open and Efficient Foundation Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.566777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.566777Z digest=sha256:a580cf1887de9fc53ae8feee537107ec07ce30b23d0e3a5cc814f2de8d6c3279

Observation 2db79da7-9150-4c5c-b483-19a604299ca4 · outbound

This paper cites On the Optimal Memorization Capacity of Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions On the Optimal Memorization Capacity of Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.197868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.197868Z digest=sha256:ab31ed6ab1093e85c608da22bd7b12fc1e139500118d50839a212d7b895924f1

Observation 83bfdf52-d4cc-4c50-bed9-d0b40e8dba71 · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.090765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.090765Z digest=sha256:563d7d3147ec12fd2e2cf0524a28086f5832012a1edb3e6364b3f6819cd9bc22

Observation 51d5d997-6339-4246-adba-6ef2c4f6b0e1 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.215985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.215985Z digest=sha256:d95a8ceccc884f6f452a88cb3afb2beb979089f582460fbabf90f3b6c7c59cf7

Pith citing papers

Observation 11a05383-c11a-49d3-8e5a-321ce04f849a · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:10.558525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:10.558525Z digest=sha256:1a1bcb0bbd55544877be5f74d4a7143c42c39360ea925b2c1fdc119db53c85fc