KinemaForge jointly infers part geometry, joint topology, and parameters from RGB-D sequences using a kinematic graph and differentiable dynamics, then verifies with an energy residual loss, reporting lower joint errors and reduced simulation drift than PARIS and Ditto baselines.
Maris: Marine open-vocabulary in- stance segmentation with geometric enhancement and se- mantic alignment
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
years
2026 10representative citing papers
Introduces MTRS task, MTRefSeg-21K benchmark of 21K image-text-mask triplets, and MTRefSeg-R1 LVLM baseline that outperforms standard models via two-stage change-aware training.
DetAS-X uses an MLLM agent to adaptively compose detection workflows from restoration modules and expert detectors, enhanced by self-evolving experience harvesting, achieving substantial F1 score gains on challenging benchmarks.
IDCL adds density-based curriculum learning and density-core guidance to deep image clustering, claiming superior robustness, faster convergence, and flexibility on benchmark datasets.
EAGC mitigates gradient entanglement in GCD by anchoring supervised gradients and adaptively projecting unlabeled ones, boosting existing methods to new state-of-the-art performance.
PALM improves long-horizon robotic manipulation success by distilling affordance representations for object interaction and predicting within-subtask progress in a VLA model.
OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.
AtmoFuseNet fuses multi-view sky cameras, millimeter-wave radar, and ceilometer data via hierarchical cross-attention, variational refinement, and motion estimation to produce 4D cloud microphysical fields and wind with reported MAEs of 0.026 g m^{-3} LWC and 1.18 m s^{-1} wind speed.
Gradient boosting with conformal prediction and mutual-information stability selection yields NAFLD risk predictions with 91.3% empirical coverage at 90% nominal level and AUROC 0.91 on multicenter Chinese data.
The NTIRE 2026 CD-FSOD Challenge report details innovative methods and performance results from 19 teams on cross-domain few-shot object detection in open- and closed-source tracks.
citing papers explorer
-
URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification
KinemaForge jointly infers part geometry, joint topology, and parameters from RGB-D sequences using a kinematic graph and differentiable dynamics, then verifies with an energy residual loss, reporting lower joint errors and reduced simulation drift than PARIS and Ditto baselines.
-
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
Introduces MTRS task, MTRefSeg-21K benchmark of 21K image-text-mask triplets, and MTRefSeg-R1 LVLM baseline that outperforms standard models via two-stage change-aware training.
-
Detect in Any Scene: An Agentic Framework for Object Detection with Experience-Aware Reasoning
DetAS-X uses an MLLM agent to adaptively compose detection workflows from restoration modules and expert detectors, enhanced by self-evolving experience harvesting, achieving substantial F1 score gains on challenging benchmarks.
-
Deep Image Clustering Based on Curriculum Learning and Density Information
IDCL adds density-based curriculum learning and density-core guidance to deep image clustering, claiming superior robustness, faster convergence, and flexibility on benchmark datasets.
-
The Devil Is in Gradient Entanglement: Energy-Aware Gradient Coordinator for Robust Generalized Category Discovery
EAGC mitigates gradient entanglement in GCD by anchoring supervised gradients and adaptively projecting unlabeled ones, boosting existing methods to new state-of-the-art performance.
-
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation
PALM improves long-horizon robotic manipulation success by distilling affordance representations for object interaction and predicting within-subtask progress in a VLA model.
-
Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline
OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.
-
Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation
AtmoFuseNet fuses multi-view sky cameras, millimeter-wave radar, and ceilometer data via hierarchical cross-attention, variational refinement, and motion estimation to produce 4D cloud microphysical fields and wind with reported MAEs of 0.026 g m^{-3} LWC and 1.18 m s^{-1} wind speed.
-
Conformal Risk Prediction for Non-Alcoholic Fatty Liver Disease Using Gradient Boosting with Distribution-Free Coverages
Gradient boosting with conformal prediction and mutual-information stability selection yields NAFLD risk predictions with 91.3% empirical coverage at 90% nominal level and AUROC 0.91 on multicenter Chinese data.
-
The Second Challenge on Cross-Domain Few-Shot Object Detection at NTIRE 2026: Methods and Results
The NTIRE 2026 CD-FSOD Challenge report details innovative methods and performance results from 19 teams on cross-domain few-shot object detection in open- and closed-source tracks.