REVIEW 6 cited by
Sequential Modeling Enables Scalable Learning for Large Vision Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
We introduce a novel sequential modeling approach which enables learning a Large Vision Model (LVM) without making use of any linguistic data. To do this, we define a common format, "visual sentences", in which we can represent raw images and videos as well as annotated data sources such as semantic segmentations and depth reconstructions without needing any meta-knowledge beyond the pixels. Once this wide variety of visual data (comprising 420 billion tokens) is represented as sequences, the model can be trained to minimize a cross-entropy loss for next token prediction. By training across various scales of model architecture and data diversity, we provide empirical evidence that our models scale effectively. Many different vision tasks can be solved by designing suitable visual prompts at test time.
Forward citations
Cited by 6 Pith papers
-
MentalThink: Shaping Thoughts in Mental SVG World
MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.
-
World Modeling with Probabilistic Structure Integration
A single probabilistic video model extracts optical flow, depth, and segments via counterfactual prompts, then integrates those structures as new token types to improve its own video predictions.
-
Is Visual in-Context Learning for Compositional Medical Tasks within Reach?
Training on synthetic compositional task sequences with sequence-level masking lets a transformer-based in-context learner follow multi-step medical imaging instructions on held-out images, but well below codebook upp...
-
CONCORD: Concept-Informed Diffusion for Dataset Distillation
A training-free concept-informed diffusion method, using LLM-retrieved and CLIP-filtered visual descriptions, improves dataset distillation accuracy on ImageNet subsets and ImageNet-1K.
-
Controllable Skin Synthesis via Lesion-Focused Vector Autoregression Model
A lesion-measurement-conditioned VAR model with a lesion-focused VQVAE achieves best average FID (0.74) for controllable skin lesion synthesis.
-
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning
A single conditional generative model, trained only on RPM-style puzzles, can be repurposed via probability scoring to solve odd-one-out, analogy, and categorization tasks, with modest zero-shot transfer.
Discussion (0). Sign in to comment.