Presents the first large-scale infrared off-road dataset and a flow-free temporal model achieving state-of-the-art freespace detection performance with real-time inference.
hub
Segment anything
23 Pith papers cite this work. Polarity classification is still indexing.
hub tools
citation-role summary
citation-polarity summary
roles
background 1polarities
background 1representative citing papers
D³ETOR combines debate-enhanced pseudo labeling from SAM with frequency-aware progressive debiasing in FADeNet to achieve state-of-the-art weakly-supervised camouflaged object detection using scribbles.
STEMbot climbs 7–33 mm stems with geometric PIN-SLAM, semantic OcTree mapping, and manifold-constrained A* planning, achieving sub-centimeter reconstructions and autonomous navigation on four plants.
SemDynReg constructs per-object ID maps from SAM and image features to regularize position, scale, and rotation of top-k Gaussians per object in dynamic 3DGS.
RGB-only active 3D scene graph generation unifies perception and planning to achieve depth-baseline parity and more than double object detection in active indoor exploration.
Fixed external cameras as Common Prior Maps boost initial object recall in 3D scene graph generation by up to 79% and improve active exploration efficiency.
MarkIt converts videos into query-conditioned marked versions via a linguistic-parsing and open-vocabulary segmentation bridge that embeds instance masks, semantic markers, and frame indices to improve Vid-LLM temporal grounding.
A variance-aware conditional MLP operating on 3D Gaussians corrects semantic errors from multi-view inconsistent 2D features to produce more accurate and robust 3D semantic Gaussian Splatting.
A new reliability score computed from the IoU difference between class-specific and class-agnostic heatmaps, boosted by adversarial enhancement, detects false negatives in binary industrial defect detectors with up to 100% recall.
AirFM-DDA reparameterizes wireless channel data into the delay-Doppler-angle domain and uses efficient window attention to achieve better zero-shot performance on channel prediction and estimation with lower compute cost.
VLMaterial fuses VLMs and physics-based radar analysis via PRCA extraction and context-augmented generation to reach 96.08% material identification accuracy on 41 everyday objects without task-specific training.
GA-GS uses motion segmentation, diffusion-based inpainting for pseudo-ground-truth, and per-Gaussian authenticity scalars to achieve SOTA static scene reconstruction from videos with dynamic occlusions.
OBEYED-VLA improves VLA robustness in cluttered real-world manipulation by disentangling perception into VLM-based object-centric grounding and geometry-aware stages, then fine-tuning the policy only on single-object demonstrations.
SolarCHIP contrastively pretrains CNN and Vision Transformer backbones on SDO AIA-HMI data with multi-granularity objectives, achieving SOTA on cross-modal translation and flare classification especially in low-resource settings.
Ground4D reconstructs dynamic 4D scenes from monocular video by initializing dynamic Gaussians from VGGT geometry and refining them with multi-view depth consistency at observed and virtual viewpoints.
SceneGraphGrounder builds a persistent 3D scene graph from VLM-inferred relations in 2D views and solves grounding via constrained graph alignment, achieving competitive zero-shot results on ScanRefer with only RGB-D input.
HiSem adds bidirectional differential attention and a two-level hierarchical routing module with MoE to handle semantic granularity differences in remote sensing change captioning, reporting +7.52% BLEU-4 on WHU-CDC.
Agentic AI systems are required to overcome the parameter coverage ceiling that prevents foundation models from handling certain out-of-distribution cases.
TopoMamba improves medical image segmentation by combining topology-aware diagonal scans with standard cross-scans and a HSIC Gate for efficient fusion, yielding gains on thin and curved targets like the pancreas.
PokeVLA is a lightweight VLA model pre-trained on 2.4M samples for spatial grounding and reasoning, then adapted via multi-view semantics and geometry alignment to achieve state-of-the-art robot manipulation performance.
Adapting Depth Anything V2 with DV-LORA bridges the ex-vivo to in-vivo gap in monocular depth estimation for specular surgical environments, achieving SOTA on SCARED and superior results on new ROCAL-T 90 dataset.
PLAF introduces a 2D pixel-wise language-aligned feature extractor paired with a redundancy-reducing storage scheme that supports accurate open-vocabulary 3D scene understanding.
A literature review of intelligent automation approaches using robotics, AI, and control for disassembly, inspection, sorting, and reprocessing of end-of-life electronics.
citing papers explorer
-
Towards All-Day Perception for Off-Road Driving: A Large-Scale Multispectral Dataset and Comprehensive Benchmark
Presents the first large-scale infrared off-road dataset and a flow-free temporal model achieving state-of-the-art freespace detection performance with real-time inference.
-
Debate-Enhanced Pseudo Labeling and Frequency-Aware Progressive Debiasing for Weakly-Supervised Camouflaged Object Detection with Scribble Annotations
D³ETOR combines debate-enhanced pseudo labeling from SAM with frequency-aware progressive debiasing in FADeNet to achieve state-of-the-art weakly-supervised camouflaged object detection using scribbles.
-
STEMbot: A Compliant Robot for Under-Canopy Plant Navigation
STEMbot climbs 7–33 mm stems with geometric PIN-SLAM, semantic OcTree mapping, and manifold-constrained A* planning, achieving sub-centimeter reconstructions and autonomous navigation on four plants.
-
SemDynReg: Semantics-Guided Deformation Regularization for Dynamic 3D Gaussian Splatting
SemDynReg constructs per-object ID maps from SAM and image features to regularize position, scale, and rotation of top-k Gaussians per object in dynamic 3DGS.
-
RGB-only Active 3D Scene Graph Generation for Indoor Mobile Robots
RGB-only active 3D scene graph generation unifies perception and planning to achieve depth-baseline parity and more than double object detection in active indoor exploration.
-
Fixed External Cameras as Common Prior Maps for Active 3D Scene Graph Generation
Fixed external cameras as Common Prior Maps boost initial object recall in 3D scene graph generation by up to 79% and improve active exploration efficiency.
-
MarkIt: Training-Free Visual Markers for Precise Video Temporal Grounding
MarkIt converts videos into query-conditioned marked versions via a linguistic-parsing and open-vocabulary segmentation bridge that embeds instance masks, semantic markers, and frame indices to improve Vid-LLM temporal grounding.
-
NRGS: Neural Regularization for Robust 3D Semantic Gaussian Splatting
A variance-aware conditional MLP operating on 3D Gaussians corrects semantic errors from multi-view inconsistent 2D features to produce more accurate and robust 3D semantic Gaussian Splatting.
-
When Can We Trust Deep Neural Networks? Towards Reliable Industrial Deployment with an Interpretability Guide
A new reliability score computed from the IoU difference between class-specific and class-agnostic heatmaps, boosted by adversarial enhancement, detects false negatives in binary industrial defect detectors with up to 100% recall.
-
AirFM-DDA: Air-Interface Foundation Model in the Delay-Doppler-Angle Domain for AI-Native 6G
AirFM-DDA reparameterizes wireless channel data into the delay-Doppler-angle domain and uses efficient window attention to achieve better zero-shot performance on channel prediction and estimation with lower compute cost.
-
VLMaterial: Vision-Language Model-Based Camera-Radar Fusion for Physics-Grounded Material Identification
VLMaterial fuses VLMs and physics-based radar analysis via PRCA extraction and context-augmented generation to reach 96.08% material identification accuracy on 41 everyday objects without task-specific training.
-
GA-GS: Generation-Assisted Gaussian Splatting for Static Scene Reconstruction
GA-GS uses motion segmentation, diffusion-based inpainting for pseudo-ground-truth, and per-Gaussian authenticity scalars to achieve SOTA static scene reconstruction from videos with dynamic occlusions.
-
Clutter-Robust Vision-Language-Action Models through Object-Centric and Geometry Grounding
OBEYED-VLA improves VLA robustness in cluttered real-world manipulation by disentangling perception into VLM-based object-centric grounding and geometry-aware stages, then fine-tuning the policy only on single-object demonstrations.
-
Contrastive Heliophysical Image Pretraining for Solar Dynamics Observatory Records
SolarCHIP contrastively pretrains CNN and Vision Transformer backbones on SDO AIA-HMI data with multi-granularity objectives, achieving SOTA on cross-modal translation and flare classification especially in low-resource settings.
-
Ground4D: Consistency-Aware 4D Reconstruction from Monocular Video
Ground4D reconstructs dynamic 4D scenes from monocular video by initializing dynamic Gaussians from VGGT geometry and refining them with multi-view depth consistency at observed and virtual viewpoints.
-
SceneGraphGrounder: Zero-Shot 3D Visual Grounding via Structured Scene Graph Matching
SceneGraphGrounder builds a persistent 3D scene graph from VLM-inferred relations in 2D views and solves grounding via constrained graph alignment, achieving competitive zero-shot results on ScanRefer with only RGB-D input.
-
HiSem: Hierarchical Semantic Disentangling for Remote Sensing Image Change Captioning
HiSem adds bidirectional differential attention and a two-level hierarchical routing module with MoE to handle semantic granularity differences in remote sensing change captioning, reporting +7.52% BLEU-4 on WHU-CDC.
-
Agentic AIs Are the Missing Paradigm for Out-of-Distribution Generalization in Foundation Models
Agentic AI systems are required to overcome the parameter coverage ceiling that prevents foundation models from handling certain out-of-distribution cases.
-
TopoMamba: Topology-Aware Scanning and Fusion for Segmenting Heterogeneous Medical Visual Media
TopoMamba improves medical image segmentation by combining topology-aware diagonal scans with standard cross-scans and a HSIC Gate for efficient fusion, yielding gains on thin and curved targets like the pancreas.
-
PokeVLA: Empowering Pocket-Sized Vision-Language-Action Model with Comprehensive World Knowledge Guidance
PokeVLA is a lightweight VLA model pre-trained on 2.4M samples for spatial grounding and reasoning, then adapted via multi-view semantics and geometry alignment to achieve state-of-the-art robot manipulation performance.
-
Bridging the Ex-Vivo to In-Vivo Gap: Synthetic Priors for Monocular Depth Estimation in Specular Surgical Environments
Adapting Depth Anything V2 with DV-LORA bridges the ex-vivo to in-vivo gap in monocular depth estimation for specular surgical environments, achieving SOTA on SCARED and superior results on new ROCAL-T 90 dataset.
-
PLAF: Pixel-wise Language-Aligned Feature Extraction for Efficient 3D Scene Understanding
PLAF introduces a 2D pixel-wise language-aligned feature extractor paired with a redundancy-reducing storage scheme that supports accurate open-vocabulary 3D scene understanding.
-
Redefining End-of-Life: Intelligent Automation for Electronics Remanufacturing Systems
A literature review of intelligent automation approaches using robotics, AI, and control for disassembly, inspection, sorting, and reprocessing of end-of-life electronics.