Pith. sign in

Paper Citation Record · LEDGER

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

As of 20 August 2026, this Paper Citation Record lists 28 of 28 outbound references and 1 inbound Pith citation observation for arXiv:2505.19531.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19531 v1

Coverage vector

measured 28 of 28 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:21:02.060803Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:57:10.558525Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

28 of 28 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfeadc09-5c90-4514-a781-34d65330052a · outbound

This paper cites Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Fast RoPE Attention: Combining the Polynomial Method and Fast Fourier Transform

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.129826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.129826Z digest=sha256:3f393cc25e3f30fa783c3e366bd4ec39b94c7aa0f9b29bef3c97eec824a60e47

Observation 3613ede7-6236-4cd7-b96b-3bef5ecbc367 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.303448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.303448Z digest=sha256:de3088d52cb8eda17b7c707479f1fd16d4480dc6201ff83a7fb384f529e6c4fe

Observation 048926cc-308b-46de-867a-170a7ef15673 · outbound

This paper cites Circuit Complexity Bounds for RoPE-based Transformer Architecture.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Circuit Complexity Bounds for RoPE-based Transformer Architecture

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:03.091343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:00.473856Z digest=sha256:0a35fe7b8297f3a7b7767361a234c2bf5659da4c215d14d1bec6f949e8de7e1b

Observation b244969a-2e56-4356-b7ec-223e8ae722a6 · outbound

This paper cites Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Provable Failure of Language Models in Learning Majority Boolean Logic via Gradient Descent

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.912542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:00.530502Z digest=sha256:2792abc60a0ea2029e99ed1982f6fe9357f7c1b31f28e7a93e5363d52b9c51da

Observation d2393cb8-fdb7-4118-bac0-38c4439db750 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.641603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.641603Z digest=sha256:0fee12379063e4b73dbdd78543886d9bae60c8b4bee14d12417fa62defab56e3

Observation 98c36a3c-2de1-4115-adcb-e61fcb878086 · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.809131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.809131Z digest=sha256:fb06affbd33cf8269a67d244da5cbf9d2a4c2da0e5cd3b41a9bac9e7512ac261

Observation 063b1293-4e99-4560-8953-4c1d2dd07be4 · outbound

This paper cites Subquadratic Algorithms and Hardness for Attention with Any Temperature.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Subquadratic Algorithms and Hardness for Attention with Any Temperature

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.673064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:00.874738Z digest=sha256:9960440ca53c63b535054aada1f57b3f6b39d5d86a4a5ab91619eec147e223e9

Observation 56df5c5a-39b4-4178-b882-02195e1018e8 · outbound

This paper cites In-Context Convergence of Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions In-Context Convergence of Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.978836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.978836Z digest=sha256:0ff62057cd0293cb312440b2a0d8865f0b570c19b881652dba44ff8caf80c49c

Observation f226550d-9ed6-48ea-858b-4d98beffd732 · outbound

This paper cites Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers Learn Nonlinear Features In Context: Nonconvex Mean-field Dynamics on the Attention Landscape

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.293029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.293029Z digest=sha256:f944f6cf0fe7b68b824e4ee7a16055e56f9d2503edb109721083dca8963df0af

Observation e4caa9f6-825f-4547-9a72-b008be475c2d · outbound

This paper cites Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Theoretical Constraints on the Expressive Power of $\mathsf{RoPE}$-based Tensor Attention Transformers

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:21:02.435527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:01.340668Z digest=sha256:35859285ce1825ef0e6a5386c77d00cab0c11f1bb021bf739c38446696e92647

Observation 15190e46-7116-43bb-bada-dbb6eb711a62 · outbound

This paper cites On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions On the Computational Capability of Graph Neural Networks: A Circuit Complexity Bound Perspective

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.419305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.419305Z digest=sha256:0a3e68224a285ff7e2f28c9e6e1c8424e05db999b0864b059e3823ca0baa255f

Observation 1027cf28-3136-40e0-9c9b-44a8bdd5b271 · outbound

This paper cites One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions One Step of Gradient Descent is Provably the Optimal In-Context Learner with One Layer of Linear Self-Attention

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.467347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.467347Z digest=sha256:200e51cff21c354943e4ec91d780a21373d4fd173c3b3eb1f927bb472dfa853c

Observation 9cadd4b5-c628-4d76-9b40-56a422be704e · outbound

This paper cites Evalu- ating numeracy of language models as a natural language inference task.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Evalu- ating numeracy of language models as a natural language inference task

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:21:03.534648Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:01.519168Z digest=sha256:b4b789780397db0b8caa6c3ab03296707b7d09c186f04ea4f71cb928c9fb7697

Observation a5275ba0-4a72-438a-b059-5c7f92ff82ec · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.699323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.699323Z digest=sha256:2e13343fcd1355c1cfb9b0fbed2ec239bd7ffaaf4c70d5f26f6ffce622de078c

Observation 595b73b5-7fae-49f6-895d-3b764859bc8a · outbound

This paper cites Transformers Learn Low Sensitivity Functions: Investigations and Implications.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers Learn Low Sensitivity Functions: Investigations and Implications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.743928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.743928Z digest=sha256:a918562888b3f2da68ccca8fde8eff41fe0c5f52345d9918b1ec19d9cadaa666

Observation e3256b49-7b24-4451-9d9e-d7816e3bed41 · outbound

This paper cites Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Sub-Task Decomposition Enables Learning in Sequence to Sequence Tasks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.786208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.786208Z digest=sha256:2bb878eea81e721ba5be2ec9b35a7ced8a2b69fe8b45217ea2c27417b2fd6d0c

Observation 939fe7ff-cdb3-4f8f-88fc-e807cdd348f8 · outbound

This paper cites From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions From Sparse Dependence to Sparse Attention: Unveiling How Chain-of-Thought Enhances Transformer Sample Efficiency

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.863974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.863974Z digest=sha256:10e871ec51627e7c8ee2235fb9fcdb8ddd439b456e9933f479e7039a63be799c

Observation 700e4e8a-caf4-45ef-bfb2-347f4c249a7d · outbound

This paper cites DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions DNABERT-2: Efficient Foundation Model and Benchmark For Multi-Species Genome

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.955998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.955998Z digest=sha256:26a6e0dd5694111ef0d29517a6739a7b2438f5f1b788e9f4d252bc7a935f9c90

Observation 64b3c3ce-f5d0-49a0-bfd9-f86abd0ff6ae · outbound

This paper cites Genomeocean: An efficient genome foundation model trained on large-scale metagenomic assemblies.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Genomeocean: An efficient genome foundation model trained on large-scale metagenomic assemblies

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:21:03.344204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T14:21:01.993859Z digest=sha256:9d9f21843972d3d759bf309befda32fb0f103bc3812dad1924e2353d9567f825

Observation fd78cc70-0225-422b-85bf-579607168ac3 · outbound

This paper cites DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions DNABERT-S: Pioneering Species Differentiation with Species-Aware DNA Embeddings

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:02.060803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:02.060803Z digest=sha256:7fb7d2999a785f0a35499bdfe08823cc8726acbc9cd7b5342883d10ec00c4fe9

Observation f0e4350d-26ba-4df3-b90e-ba5b1d39be14 · outbound

This paper cites Transformers in Uniform TC$^0$.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Transformers in Uniform TC$^0$

Reference 1867

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.402417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.402417Z digest=sha256:74042854c763efba9baa8537fbf0787b09bd0b608bb3113c9315d40eb787f74e

Observation 206e2655-941d-4e3c-a85b-05d3558c2bba · outbound

This paper cites Why are Sensitive Functions Hard for Transformers?.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Why are Sensitive Functions Hard for Transformers?

Reference 1963

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.050664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.050664Z digest=sha256:36d03720fd6e81158064ef45a80c2d193827a1e33932c2a75c429cc178c136cd

Observation 23287124-bca5-4cf0-8d95-87a8944e8a85 · outbound

This paper cites Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Can You Count to Nine? A Human Evaluation Benchmark for Counting Limits in Modern Text-to-Video Models

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.742146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.742146Z digest=sha256:597bdf891b1e7f09c91cd9040aa3605ced1ea2e5227bd9cb614ed6b22301a7a0

Observation 826301cb-e573-4121-bba2-e1db4ad29938 · outbound

This paper cites Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Are Transformers with One Layer Self-Attention Using Low-Rank Weight Matrices Universal Approximators?

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.122910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.122910Z digest=sha256:a80f28b8c5c75c6906170241f41974a8df2cf42b892a498ab66ee06ddd1f6bbc

Observation ba15c0a1-cce9-4037-83c4-9e7db19ab1dc · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions LLaMA: Open and Efficient Foundation Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.566777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.566777Z digest=sha256:c8594adc5c05f8c92edce5262a2c513c1c0dcd2c1f9893bc7de9031836fa5990

Observation 2db79da7-9150-4c5c-b483-19a604299ca4 · outbound

This paper cites On the Optimal Memorization Capacity of Transformers.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions On the Optimal Memorization Capacity of Transformers

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:01.197868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:01.197868Z digest=sha256:4367f7ccdadf50c3c32caae67bd03b282d124674c9e8c4a13776a9a1b6da8e26

Observation 83bfdf52-d4cc-4c50-bed9-d0b40e8dba71 · outbound

This paper cites How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions How to Capture Higher-order Correlations? Generalizing Matrix Softmax Attention to Kronecker Computation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.090765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.090765Z digest=sha256:d78ad820ae9347e5889093620828d1bf7584e2fd5770bb2a4c2600e13c323231

Observation 51d5d997-6339-4246-adba-6ef2c4f6b0e1 · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Minimalist Softmax Attention Provably Learns Constrained Boolean Functions Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T14:21:00.215985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:21:00.215985Z digest=sha256:ea9337fc218117f40e6c6ba2a72052f049611f7c9a96e05db0f86b615770841e

Pith citing papers

Observation 11a05383-c11a-49d3-8e5a-321ce04f849a · inbound

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently cites this paper.

Transformers with RL or SFT Provably Learn Sparse Boolean Functions, But Differently Minimalist Softmax Attention Provably Learns Constrained Boolean Functions

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T20:57:10.558525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T20:57:10.558525Z digest=sha256:38f151cd6bca701269439bd1e158dc288029283dfd179c32a66d1e94b4db4c20