Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2312.06528.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T06:10:17.138258Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-21T07:34:02.998394Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ecfbb607-d2bd-4d8a-bb96-9b0fc2788508 · inbound
A Theoretical Study of (Hyper) Self-Attention through the Lens of Interactions: Representation, Training, Generalization Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9de054c-1af0-4782-a4fd-5f0fc9b868b9 · inbound
Spectral Transformer Neural Processes Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d94cfa03-61e1-4974-8dd2-0104b67b1879 · inbound
One for All: A Non-Linear Transformer can Enable Cross-Domain Generalization for In-Context Reinforcement Learning Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eaefecc6-06a6-4095-9039-12c0374c2ae9 · inbound
Multi-Head Attention as Ensemble Nadaraya-Watson Estimation: Variance Reduction, Decorrelation, and Optimal Head Diversity Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e5ea04a7-25a0-4d51-a9d4-a1c7108f0e37 · inbound
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex Transformers Implement Functional Gradient Descent to Learn Non-Linear Functions In Context
Reference 161
Source-reported events for the cited work
Unavailable: canonical work link unavailable.