OctoSense supplies a large multimodal robotics dataset and a late-fusion masked autoencoder that runs fast and outperforms image-only models on optical flow, depth, segmentation, and ego-motion tasks while remaining robust under sensor degradation.
arXiv preprint arXiv:2509.25146 (2025)
2 Pith papers cite this work. Polarity classification is still indexing.
2
Pith papers citing it
fields
cs.CV 2years
2026 2verdicts
UNVERDICTED 2representative citing papers
A single event-camera feature matching model trained on synthetic data achieves zero-shot wide-baseline correspondence across unseen datasets with 37.7% improvement over prior methods.
citing papers explorer
-
OctoSense: Self-Supervised Learning for Multimodal Robot Perception
OctoSense supplies a large multimodal robotics dataset and a late-fusion masked autoencoder that runs fast and outperforms image-only models on optical flow, depth, segmentation, and ego-motion tasks while remaining robust under sensor degradation.
-
Match-Any-Events: Zero-Shot Motion-Robust Feature Matching Across Wide Baselines for Event Cameras
A single event-camera feature matching model trained on synthetic data achieves zero-shot wide-baseline correspondence across unseen datasets with 37.7% improvement over prior methods.