REVIEW 10 cited by
Foundational Models Defining a New Era in Vision: A Survey and Outlook
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Vision systems to see and reason about the compositional nature of visual scenes are fundamental to understanding our world. The complex relations between objects and their locations, ambiguities, and variations in the real-world environment can be better described in human language, naturally governed by grammatical rules and other modalities such as audio and depth. The models learned to bridge the gap between such modalities coupled with large-scale training data facilitate contextual reasoning, generalization, and prompt capabilities at test time. These models are referred to as foundational models. The output of such models can be modified through human-provided prompts without retraining, e.g., segmenting a particular object by providing a bounding box, having interactive dialogues by asking questions about an image or video scene or manipulating the robot's behavior through language instructions. In this survey, we provide a comprehensive review of such emerging foundational models, including typical architecture designs to combine different modalities (vision, text, audio, etc), training objectives (contrastive, generative), pre-training datasets, fine-tuning mechanisms, and the common prompting patterns; textual, visual, and heterogeneous. We discuss the open challenges and research directions for foundational models in computer vision, including difficulties in their evaluations and benchmarking, gaps in their real-world understanding, limitations of their contextual understanding, biases, vulnerability to adversarial attacks, and interpretability issues. We review recent developments in this field, covering a wide range of applications of foundation models systematically and comprehensively. A comprehensive list of foundational models studied in this work is available at \url{https://github.com/awaisrauf/Awesome-CV-Foundational-Models}.
Forward citations
Cited by 10 Pith papers
-
GHOST: Geometry-Guided Hallucination of Opaque Surface Textures
Geometry-guided hallucination of opaque textures from transparent regions lets off-the-shelf depth and reconstruction models recover accurate surfaces without retraining.
-
Your Spending Needs Attention: Modeling Financial Habits with Transformers
A causal transformer pre-trained with next-token prediction on tokenized bank transactions, fused end-to-end with tabular features, lifts recommendation test AUC by 1.25% relative over a LightGBM baseline at Nubank.
-
M-SpecGene: Generalized Foundation Model for RGBT Multispectral Vision
M-SpecGene is a Siamese masked-autoencoder foundation model for RGB-thermal vision, trained on the RGBT550K dataset with a GMM-CMSS progressive masking strategy, and evaluated on four downstream tasks.
-
Manifold-Constrained Hyper-Connections for Parameter-Efficient Finetuning
Applying mHC as a PEFT method shows that learned residual mixing is unnecessary — even harmful — in finetuning, and mHC+LoRA combinations give small task-dependent gains.
-
A Modality-agnostic Multi-task Foundation Model for Human Brain Imaging
A multi-task model trained on synthetic plus real brain scans performs synthesis, segmentation, registration, distance-map prediction, and bias-field estimation across T1w, T2w, FLAIR MRI and CT without fine-tuning.
-
Context-Adaptive Inference: A Unified Statistical and Foundation-Model View
Under linear, squared-loss assumptions, explicit context adaptation and in-context learning both reduce to kernel ridge regression on joint input-context features.
-
Towards Affordable Tumor Segmentation and Visualization for 3D Breast MRI Using SAM2
SAM2 can segment breast tumors in 3D MRI with a single bounding-box prompt, and center-outward propagation yields the best volumetric Dice.
-
Foundation Models for Astrophysics
Astronomical 'foundation models' largely reuse transformers and self-supervised pretraining, but evidence of transfer to new instruments, populations, or tasks remains rare; the paper argues such evidence, not archite...
-
Scout: Leveraging Large Language Models for Rapid Digital Evidence Discovery
Scout applies off-the-shelf LLMs and vision models to triage digital evidence, but only anecdotal examples are shown and accuracy is withheld.
-
Vision Generalist Model: A Survey
A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.
Discussion (0). Sign in to comment.