Pith. sign in

REVIEW 2 cited by

Building Blocks for a Complex-Valued Transformer Architecture

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.09827 v1 pith:FN2YI3FY submitted 2023-06-16 cs.LG cs.CVcs.NE

classification cs.LGcs.CVcs.NE
keywords complex-valuedsignalsarchitecturereal-valuedtransformerapplicationsblocksbuilding
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Most deep learning pipelines are built on real-valued operations to deal with real-valued inputs such as images, speech or music signals. However, a lot of applications naturally make use of complex-valued signals or images, such as MRI or remote sensing. Additionally the Fourier transform of signals is complex-valued and has numerous applications. We aim to make deep learning directly applicable to these complex-valued signals without using projections into $\mathbb{R}^2$. Thus we add to the recent developments of complex-valued neural networks by presenting building blocks to transfer the transformer architecture to the complex domain. We present multiple versions of a complex-valued Scaled Dot-Product Attention mechanism as well as a complex-valued layer normalization. We test on a classification and a sequence generation task on the MusicNet dataset and show improved robustness to overfitting while maintaining on-par performance when compared to the real-valued transformer architecture.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Complex-Valued Phase-Coherent Transformer

    cs.LG 2026-05 unverdicted novelty 7.0 of 10 partial

    Sigmoid gating on L2-normalised complex cosine scores, with no row normalisation, generalises across long-range, positional, phase and vision tasks, though the depth-stability theorem assumes its own substance.

  2. IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Self-supervised latent prediction on raw complex ultrasound channel data reduces the labeled data needed for sound-speed estimation by roughly 3-4x in simulation, reaching 15.6 m/s error with 10,000 labels.

Pith tools