Pano3D augments 3D feedforward reconstruction backbones with a set-based mask decoder and joint geometric-semantic training to achieve SOTA 3D panoptic segmentation on ScanNet, ScanNet200, and ScanNet++.
hub
In: ICML (2021)
10 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
fields
cs.CV 10roles
method 1polarities
use method 1representative citing papers
Topo-R1 fine-tunes a vision-language model using a topology-aware reward and GRPO to detect anomalies such as broken or spurious connections in tubular segmentation masks, outperforming standard VLMs.
SCENT uses VLM-generated scene descriptions as a semantic bridge to align electronic-nose signals with visual and textual embeddings, improving cross-modal smell retrieval and enabling object-context odor disentanglement.
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
Pretrained vision transformers exhibit strong intra-object leakage where each part representation encodes information from the entire object, undermining the faithfulness of attention-based part-centric interpretability methods.
HO-Flow synthesizes realistic hand-object motions from text and canonical 3D objects via an interaction-aware VAE and masked flow matching, reporting SOTA physical plausibility and diversity on GRAB, OakInk, and DexYCB.
MM1 models achieve state-of-the-art few-shot multimodal results by pre-training on a careful mix of image-caption, interleaved, and text-only data with optimized image encoders.
A new orientation-aware motion encoding network combined with adapted text prompts improves zero-shot action recognition across NTU-RGB+D, BABEL, NW-UCLA and surveillance datasets by addressing viewpoint domain gaps.
LATERN reformulates video anomaly detection as temporal evidence aggregation via context-aware scoring (CEA) and recursive aggregation (REA) to improve accuracy and coherence for frozen VLMs on benchmarks like UCF-Crime.
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.
citing papers explorer
-
Pano3D: Unified 3D Reconstruction and Panoptic Segmentation
Pano3D augments 3D feedforward reconstruction backbones with a set-based mask decoder and joint geometric-semantic training to achieve SOTA 3D panoptic segmentation on ScanNet, ScanNet200, and ScanNet++.
-
Topo-R1: Detecting Topological Anomalies via Vision-Language Models
Topo-R1 fine-tunes a vision-language model using a topology-aware reward and GRPO to detect anomalies such as broken or spurious connections in tubular segmentation masks, outperforming standard VLMs.
-
What Images Cannot Say: Language-Guided Olfactory Representation Learning
SCENT uses VLM-generated scene descriptions as a semantic bridge to align electronic-nose signals with visual and textual embeddings, improving cross-modal smell retrieval and enabling object-context odor disentanglement.
-
FlowCIR: Semantic Transport via Flow Matching for Zero-Shot Composed Image Retrieval
FlowCIR frames ZS-CIR as conditional flow matching transport on fixed VLM embeddings plus an inference-time Multi-Negative Steering fix for negation, reporting competitive benchmark results at far lower training cost.
-
Metonymy in vision models undermines attention-based interpretability
Pretrained vision transformers exhibit strong intra-object leakage where each part representation encodes information from the entire object, undermining the faithfulness of attention-based part-centric interpretability methods.
-
HO-Flow: Generalizable Hand-Object Interaction Generation with Latent Flow Matching
HO-Flow synthesizes realistic hand-object motions from text and canonical 3D objects via an interaction-aware VAE and masked flow matching, reporting SOTA physical plausibility and diversity on GRAB, OakInk, and DexYCB.
-
MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training
MM1 models achieve state-of-the-art few-shot multimodal results by pre-training on a careful mix of image-caption, interleaved, and text-only data with optimized image encoders.
-
Cross-Domain Human Action Recognition from Multiview Motion and Textual Descriptions
A new orientation-aware motion encoding network combined with adapted text prompts improves zero-shot action recognition across NTU-RGB+D, BABEL, NW-UCLA and surveillance datasets by addressing viewpoint domain gaps.
-
LATERN: Test-Time Context-Aware Explainable Video Anomaly Detection
LATERN reformulates video anomaly detection as temporal evidence aggregation via context-aware scoring (CEA) and recursive aggregation (REA) to improve accuracy and coherence for frozen VLMs on benchmarks like UCF-Crime.
-
Accelerated Likelihood Maximization for Diffusion-based Versatile Content Generation
ALM integrates likelihood maximization and acceleration into diffusion reverse sampling to enable globally coherent generation from incomplete inputs.