A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
Fashionpedia: Ontology, segmentation, and an attribute localization dataset
14 Pith papers cite this work, alongside 48 external citations. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
background 2polarities
background 2representative citing papers
Introduces LSM that outputs calibrated multimodal spatial distributions from language plus scene graph, fused via VL-Map to improve 3D target localization on VLA-3D benchmark and real robot.
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
Point cloud geometry is cast as a statistical manifold of per-point Gaussians, with POLI learning the mapping self-supervisedly to improve perception without labeled data.
CAAT selects critical parameters for adversarial robustness in ViTs and applies PEFT to tune only those, yielding a 4.3% robustness drop versus full AT while using ~6% of parameters.
Holi-DETR improves fashion item detection by integrating co-occurrence probabilities, inter-item spatial arrangements, and body keypoint relationships into the DETR architecture.
A 2D Gaussian Splatting method with depth map generation and divide-and-conquer strategy produces high-quality TDOMs and spatial reconstructions without explicit DSM or occlusion detection.
Proprioceptive-visual correspondence lets a humanoid robot acquire self-other distinction and a 3D self-model without labels or kinematics.
LUSIS-DETR with AquaBSAM reports leading performance on four underwater instance segmentation datasets and real-time FP16 inference on an NVIDIA T4 GPU.
Model interpretation methods are reformulated to emphasize baselines; gradient-based methods, IG, and Taylor expansion are unified with explicit baselines identified, and a revised IG is developed for improved results from any layer.
Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.
N2 injection achieves stationary ELM-free H-mode in EAST full-metal-wall tokamak via pedestal-foot DTEM regulation of edge gradients.
A decoupled prototype matching approach with vision foundation models delivers 6.9% higher average precision than prior training-free methods on industrial few-shot object detection benchmarks.
citing papers explorer
-
Vision as Unified Multimodal Generation
A single unified multimodal model matches leading task-specialized vision systems across detection, segmentation, dense geometry, and multi-view 3D by casting all outputs as native text or image generation.
-
Language as a Sensor: Calibrated Spatial Belief Estimation in 3D Scenes from Natural Language
Introduces LSM that outputs calibrated multimodal spatial distributions from language plus scene graph, fused via VL-Map to improve 3D target localization on VLA-3D benchmark and real robot.
-
SegRAG: Training-Free Retrieval-Augmented Semantic Segmentation
SegRAG is a training-free retrieval-augmented framework that extracts class-specific point prompts from a filtered DINOv3 feature bank to boost SAM3 semantic segmentation performance on standard and agricultural benchmarks.
-
Learning Point Cloud Geometry as a Statistical Manifold: Theory and Practice
Point cloud geometry is cast as a statistical manifold of per-point Gaussians, with POLI learning the mapping self-supervisedly to improve perception without labeled data.
-
Efficient Adversarial Training via Criticality-Aware Fine-Tuning
CAAT selects critical parameters for adversarial robustness in ViTs and applies PEFT to tune only those, yielding a 4.3% robustness drop versus full AT while using ~6% of parameters.
-
Holi-DETR: Holistic Fashion Item Detection Leveraging Contextual Information
Holi-DETR improves fashion item detection by integrating co-occurrence probabilities, inter-item spatial arrangements, and body keypoint relationships into the DETR architecture.
-
High-Quality Spatial Reconstruction and Orthoimage Generation Using Efficient 2D Gaussian Splatting
A 2D Gaussian Splatting method with depth map generation and divide-and-conquer strategy produces high-quality TDOMs and spatial reconstructions without explicit DSM or occlusion detection.
-
Proprioceptive-visual correspondence enables self-other distinction in humanoid robots
Proprioceptive-visual correspondence lets a humanoid robot acquire self-other distinction and a 3D self-model without labels or kinematics.
-
Aqua Boundary-Saliency Attention Module for Lightweight Underwater Salient Instance Segmentation Detection Transformer
LUSIS-DETR with AquaBSAM reports leading performance on four underwater instance segmentation datasets and real-time FP16 inference on an NVIDIA T4 GPU.
-
The Neglected Baseline in Model Interpretation
Model interpretation methods are reformulated to emphasize baselines; gradient-based methods, IG, and Taylor expansion are unified with explicit baselines identified, and a revised IG is developed for improved results from any layer.
-
Efficient 3D Content Reconstruction and Generation
Presents Instant3D for rapid text/image-to-3D generation via multi-view diffusion plus feed-forward reconstruction, and FastMap for 10x faster structure-from-motion with comparable accuracy.
-
Nitrogen-induced ELM suppression and confinement improvement in the EAST tokamak with a full metal wall
N2 injection achieves stationary ELM-free H-mode in EAST full-metal-wall tokamak via pedestal-foot DTEM regulation of edge gradients.
-
Decoupled Prototype Matching with Vision Foundation Models for Few-Shot Industrial Object Detection
A decoupled prototype matching approach with vision foundation models delivers 6.9% higher average precision than prior training-free methods on industrial few-shot object detection benchmarks.
- MotionMAR: Multi-scale Auto-Regressive Human Motion Reconstruction from Sparse Observations