A multimodal pipeline decodes EEG into 3D meshes via EEG-to-image, MLLM reasoning, diffusion, and single-image-to-3D conversion, reporting 85.4% 10-way accuracy and 0.648 CLIPScore.
In: 2009 IEEE conference on computer vision and pattern recognition
13 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
dataset 1polarities
use dataset 1representative citing papers
A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
TransUNet is a hybrid CNN-Transformer architecture that outperforms prior U-Net and Transformer baselines on multi-organ and cardiac medical image segmentation tasks.
Single-level feature-to-feature forecasting with deformable convolutions on coarse abstract features from a segmentation backbone achieves state-of-the-art results for nine-timestep future semantic segmentation on Cityscapes validation.
SOCS derives per-step closed-form control signals from stochastic optimal control to steer diffusion sampling trajectories toward measurements while preserving the generative prior.
DO-ALL applies dataset distillation to generate synthetic source anchors that stabilize continual test-time adaptation under evolving domains without storing original source data.
Domain-adapted augmentations and plant-specific training data improve self-supervised representations for fine-grained plant species recognition over standard SSL pipelines.
EEG2Vision reconstructs images from EEG using diffusion models plus LLM-guided boosting, with reconstruction quality holding up reasonably as electrode count drops from 128 to 24 channels.
DE-CM trains a flow-map consistency model on three sub-trajectories (coupling, instantaneous, noise-to-noisy) and reports 1.70 FID one-step on ImageNet 256.
Neural networks trained via supervised contrastive learning yield feature attributions that are more faithful, less complex, and more continuous than those from cross-entropy trained networks.
Self-supervised contrastive learning adapts ViT for cardiac MR classification, outperforming supervised training with AUC >0.75 on four common sequences and generalization to BraTS and ADNI.
citing papers explorer
-
Brain3D: EEG-to-3D Decoding of Visual Representations via Multimodal Reasoning
A multimodal pipeline decodes EEG into 3D meshes via EEG-to-image, MLLM reasoning, diffusion, and single-image-to-3D conversion, reporting 85.4% 10-way accuracy and 0.648 CLIPScore.
-
Self-supervised pretraining for an iterative image size agnostic vision transformer
A sequential-to-global SSL method based on DINO pretrains iterative foveal-inspired vision transformers to achieve competitive ImageNet-1K performance with constant compute regardless of input resolution.
-
TransUNet: Transformers Make Strong Encoders for Medical Image Segmentation
TransUNet is a hybrid CNN-Transformer architecture that outperforms prior U-Net and Transformer baselines on multi-organ and cardiac medical image segmentation tasks.
-
Single Level Feature-to-Feature Forecasting with Deformable Convolutions
Single-level feature-to-feature forecasting with deformable convolutions on coarse abstract features from a segmentation backbone achieves state-of-the-art results for nine-timestep future semantic segmentation on Cityscapes validation.
-
Stochastic Optimal Control Sampling for Diffusion Inverse Problems
SOCS derives per-step closed-form control signals from stochastic optimal control to steer diffusion sampling trajectories toward measurements while preserving the generative prior.
-
Distill Once, Adapt Life-Long: Exploring Dataset Distillation for Continual Test-Time Adaptation
DO-ALL applies dataset distillation to generate synthetic source anchors that stabilize continual test-time adaptation under evolving domains without storing original source data.
-
Self-Supervised Learning of Plant Image Representations
Domain-adapted augmentations and plant-specific training data improve self-supervised representations for fine-grained plant species recognition over standard SSL pipelines.
-
EEG2Vision: A Multimodal EEG-Based Framework for 2D Visual Reconstruction in Cognitive Neuroscience
EEG2Vision reconstructs images from EEG using diffusion models plus LLM-guided boosting, with reconstruction quality holding up reasonably as electrode count drops from 128 to 24 channels.
-
Dual-End Consistency Model
DE-CM trains a flow-map consistency model on three sub-trajectories (coupling, instantaneous, noise-to-noisy) and reports 1.70 FID one-step on ImageNet 256.
-
On the Properties of Feature Attribution for Supervised Contrastive Learning
Neural networks trained via supervised contrastive learning yield feature attributions that are more faithful, less complex, and more continuous than those from cross-entropy trained networks.
-
Self-Supervised Contrastive Learning for Cardiac MR Sequence Classification
Self-supervised contrastive learning adapts ViT for cardiac MR classification, outperforming supervised training with AUC >0.75 on four common sequences and generalization to BraTS and ADNI.
- AdaBoosting Text Prompts for Vision-Language Models
- Free-Flow Class-Incremental Learning: Towards Robust CIL under Variable Class Arrivals