Pith. sign in

REVIEW 2 cited by

MIPD: A Multi-sensory Interactive Perception Dataset for Embodied Intelligent Driving

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.05881 v1 pith:AEP6YGCJ submitted 2024-11-08 cs.RO

classification cs.RO
keywords drivingdatasetautonomousembodiedinformationmipdcurrentinputs
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

During the process of driving, humans usually rely on multiple senses to gather information and make decisions. Analogously, in order to achieve embodied intelligence in autonomous driving, it is essential to integrate multidimensional sensory information in order to facilitate interaction with the environment. However, the current multi-modal fusion sensing schemes often neglect these additional sensory inputs, hindering the realization of fully autonomous driving. This paper considers multi-sensory information and proposes a multi-modal interactive perception dataset named MIPD, enabling expanding the current autonomous driving algorithm framework, for supporting the research on embodied intelligent driving. In addition to the conventional camera, lidar, and 4D radar data, our dataset incorporates multiple sensor inputs including sound, light intensity, vibration intensity and vehicle speed to enrich the dataset comprehensiveness. Comprising 126 consecutive sequences, many exceeding twenty seconds, MIPD features over 8,500 meticulously synchronized and annotated frames. Moreover, it encompasses many challenging scenarios, covering various road and lighting conditions. The dataset has undergone thorough experimental validation, producing valuable insights for the exploration of next-generation autonomous driving frameworks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TEM^3-Learning: Time-Efficient Multimodal Multi-Task Learning for Advanced Assistive Driving

    cs.CV 2025-06 conditional novelty 4.0 of 10

    A multimodal multi-task architecture combining Mamba-based temporal-spatial features and task-specific gating achieves state-of-the-art accuracy on the AIDE assistive-driving benchmark at real-time speed.

  2. VM-BHINet:Vision Mamba Bimanual Hand Interaction Network for 3D Interacting Hand Mesh Recovery From a Single RGB Image

    cs.CV 2025-04 reject novelty 4.0 of 10

    VM-BHINet combines a Vision Mamba block with an interaction feature module to recover two interacting hand meshes from one RGB image, reporting 5.44 mm MPVPE and 5.09 mm MPJPE on InterHand2.6M.

Pith tools