GazeWorld autoregressively predicts latent representations of radiologist fixated patches with a spatial-completion branch to pretrain features that achieve SOTA supervised and zero-shot diagnostic accuracy on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax.
arXiv preprint arXiv:2401.08541 , year=
11 Pith papers cite this work, alongside 6 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
representative citing papers
Optimal INR freeze depth matches highest weight stable rank layer; SAEs reveal SIREN atoms are localized while FFMLP atoms trace cohort contours with causal impact on PSNR.
HPP decouples perception from reasoning in long-video VLMs by having an LLM run iterative programmatic probes on hierarchically segmented video, reporting gains on LongVideoBench, EgoSchema, VideoMME, and MLVU.
Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains increasing with model depth across modalities.
Weighted Reverse Convolution is a spatially adaptive inverse operator for densifying high-level visual descriptors from vision foundation models, using weighted regularization and an FFT closed-form solution to improve dense prediction tasks.
A dual-branch foundation-model embedding-difference approach with low-rank adaptation lowers average BSCER100 from 6.16% to 2.17% in cross-database known-attack D-MAD.
SmolVLA is a small efficient VLA model that achieves performance comparable to 10x larger models while training on one GPU and deploying on consumer hardware via community data and chunked asynchronous action prediction.
MM1 models achieve state-of-the-art few-shot multimodal results by pre-training on a careful mix of image-caption, interleaved, and text-only data with optimized image encoders.
Benchmark of Affine, AIM, JetFormer and VQ-VAE tokenizers on galaxy images shows decoupled reconstruction and representation performance with no consistent winner.
TaTok is a theoretically grounded adaptive tokenization method that uses global tokens and cumulative conditional entropy filtering to reduce redundancy while improving reconstruction quality over fixed-rate patch tokenization.
ReasonCLIP-58M applies continual pretraining with visually grounded reasoning captions on 58M examples to improve CLIP-style models on commonsense and compositional reasoning tasks.
citing papers explorer
-
A World Model of Radiologist Reading for Medical Image Representation Learning
GazeWorld autoregressively predicts latent representations of radiologist fixated patches with a spatial-completion branch to pretrain features that achieve SOTA supervised and zero-shot diagnostic accuracy on CheXpert, RSNA Pneumonia, and SIIM-ACR Pneumothorax.
-
What Cohort INRs Encode and Where to Freeze Them
Optimal INR freeze depth matches highest weight stable rank layer; SAEs reveal SIREN atoms are localized while FFMLP atoms trace cohort contours with causal impact on PSNR.
-
HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning
HPP decouples perception from reasoning in long-video VLMs by having an LLM run iterative programmatic probes on hierarchically segmented video, reporting gains on LongVideoBench, EgoSchema, VideoMME, and MLVU.
-
Uncovering the Latent Potential of Deep Intermediate Representations
Introduces LOES, a constructive spectral method to select task-discriminative subspaces from intermediate layer embeddings, and GeoReg for enforcing simplicial class geometry during fine-tuning, with reported gains increasing with model depth across modalities.
-
Weighted Reverse Convolution for Feature Upsampling
Weighted Reverse Convolution is a spatially adaptive inverse operator for densifying high-level visual descriptors from vision foundation models, using weighted regularization and an FFT closed-form solution to improve dense prediction tasks.
-
DifFoundMAD: Foundation Models meet Differential Morphing Attack Detection
A dual-branch foundation-model embedding-difference approach with low-rank adaptation lowers average BSCER100 from 6.16% to 2.17% in cross-database known-attack D-MAD.
-
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics
SmolVLA is a small efficient VLA model that achieves performance comparable to 10x larger models while training on one GPU and deploying on consumer hardware via community data and chunked asynchronous action prediction.
-
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
MM1 models achieve state-of-the-art few-shot multimodal results by pre-training on a careful mix of image-caption, interleaved, and text-only data with optimized image encoders.
-
The Galaxy's Guide to the Tokenizer: A Benchmark for Scientific Foundation Models
Benchmark of Affine, AIM, JetFormer and VQ-VAE tokenizers on galaxy images shows decoupled reconstruction and representation performance with no consistent winner.
-
Mutual Enhancement Between Global Tokens and Patch Tokens: From Theory to Practice
TaTok is a theoretically grounded adaptive tokenization method that uses global tokens and cumulative conditional entropy filtering to reduce redundancy while improving reconstruction quality over fixed-rate patch tokenization.
-
ReasonCLIP-58M: Visually Grounded Commonsense Reasoning Supervision for CLIP
ReasonCLIP-58M applies continual pretraining with visually grounded reasoning captions on 58M examples to improve CLIP-style models on commonsense and compositional reasoning tasks.