Vision-transformer neural networks trained on simulated charge stability diagrams from a disordered generalized Hubbard model predict SOC-induced spin-flip tunneling amplitudes with R² ≈ 0.94 even when other parameters are unknown.
Title resolution pending
9 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
PARCEL is a new visual tokenization architecture combining pool-anchored resampling with conditioned elastic queries to enhance performance-efficiency tradeoffs in LVLMs over prior matryoshka methods.
Polynomial representations yield an effective-degree simplicity metric that predicts generalization across tasks and serves as a differentiable regularizer improving performance in classification and RL.
SPT generalizes SICGAT and ViT to arbitrary superpixel chunking with multidimensional sine-cosine encodings and shape/color-enriched patches, outperforming prior GNN superpixel methods while matching ViTs on CIFAR10, FashionMNIST, and Imagenette.
Fine-tuned image classifiers, led by RegNetY-16GF at about 99% top-1 accuracy, can sort historical archaeological page scans into 11 content categories, and the authors release the dataset, code, and model weights.
Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.
Decoupled weight decay proportional to gamma squared yields stable weight and gradient norms under the steady-state assumption that updates are independent of weights.
Fine-tuned RegNetY-16GF reaches 99.16% accuracy classifying 48k century-old Czech archival pages into 11 visual content types, beating a 75% hand-crafted feature baseline, with models and data released publicly.
Self-supervised contrastive learning adapts ViT for cardiac MR classification, outperforming supervised training with AUC >0.75 on four common sequences and generalization to BraTS and ADNI.
citing papers explorer
-
Predicting spin-orbit coupling in hole spin qubit arrays with vision-transformer-based neural networks on a generalized Hubbard model
Vision-transformer neural networks trained on simulated charge stability diagrams from a disordered generalized Hubbard model predict SOC-induced spin-flip tunneling amplitudes with R² ≈ 0.94 even when other parameters are unknown.
-
PARCEL: Pool-Anchored Resampling with Conditioned Elastic Queries for Efficient Vision-Language Understanding
PARCEL is a new visual tokenization architecture combining pool-anchored resampling with conditioned elastic queries to enhance performance-efficiency tradeoffs in LVLMs over prior matryoshka methods.
-
Quantifying and Optimizing Simplicity via Polynomial Representations
Polynomial representations yield an effective-degree simplicity metric that predicts generalization across tasks and serves as a differentiable regularizer improving performance in classification and RL.
-
Is an Image Also Worth 16x16=256 Superpixels? A Framework for Attentional Image Classification
SPT generalizes SICGAT and ViT to arbitrary superpixel chunking with multidimensional sine-cosine encodings and shape/color-enriched patches, outperforming prior GNN superpixel methods while matching ViTs on CIFAR10, FashionMNIST, and Imagenette.
-
Page image classification for content-specific data processing
Fine-tuned image classifiers, led by RegNetY-16GF at about 99% top-1 accuracy, can sort historical archaeological page scans into 11 content categories, and the authors release the dataset, code, and model weights.
-
Do Video Foundation Models Understand Intuitive Physics? A Layerwise Probing Analysis
Video foundation models encode intuitive physics knowledge that is strongest in V-JEPA at intermediate-to-late layers and depends on pretraining type and probe design.
-
Correction of Decoupled Weight Decay
Decoupled weight decay proportional to gamma squared yields stable weight and gradient norms under the steady-state assumption that updates are independent of weights.
-
Page image classifier fine-tuned on century-spanning archives of scanned documents for further content-specific processing
Fine-tuned RegNetY-16GF reaches 99.16% accuracy classifying 48k century-old Czech archival pages into 11 visual content types, beating a 75% hand-crafted feature baseline, with models and data released publicly.
-
Self-Supervised Contrastive Learning for Cardiac MR Sequence Classification
Self-supervised contrastive learning adapts ViT for cardiac MR classification, outperforming supervised training with AUC >0.75 on four common sequences and generalization to BraTS and ADNI.