REVIEW 2 cited by
Building Blocks for a Complex-Valued Transformer Architecture
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
Most deep learning pipelines are built on real-valued operations to deal with real-valued inputs such as images, speech or music signals. However, a lot of applications naturally make use of complex-valued signals or images, such as MRI or remote sensing. Additionally the Fourier transform of signals is complex-valued and has numerous applications. We aim to make deep learning directly applicable to these complex-valued signals without using projections into $\mathbb{R}^2$. Thus we add to the recent developments of complex-valued neural networks by presenting building blocks to transfer the transformer architecture to the complex domain. We present multiple versions of a complex-valued Scaled Dot-Product Attention mechanism as well as a complex-valued layer normalization. We test on a classification and a sequence generation task on the MusicNet dataset and show improved robustness to overfitting while maintaining on-par performance when compared to the real-valued transformer architecture.
Forward citations
Cited by 2 Pith papers
-
Complex-Valued Phase-Coherent Transformer
Sigmoid gating on L2-normalised complex cosine scores, with no row normalisation, generalises across long-range, positional, phase and vision tasks, though the depth-stability theorem assumes its own substance.
-
IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
Self-supervised latent prediction on raw complex ultrasound channel data reduces the labeled data needed for sound-speed estimation by roughly 3-4x in simulation, reaching 15.6 m/s error with 10,000 labels.
Discussion (0). Continue with ORCID to comment.