VER is a multimodal-LLM framework that extracts low-dimensional coordinates from video and, together with SINDy, discovers governing ODEs without fine-tuning the LLM.
Vision-based Discovery of Nonlinear Dynamics for 3D Moving Target
1 Pith paper cite this work. Polarity classification is still indexing.
abstract
Data-driven discovery of governing equations has kindled significant interests in many science and engineering areas. Existing studies primarily focus on uncovering equations that govern nonlinear dynamics based on direct measurement of the system states (e.g., trajectories). Limited efforts have been placed on distilling governing laws of dynamics directly from videos for moving targets in a 3D space. To this end, we propose a vision-based approach to automatically uncover governing equations of nonlinear dynamics for 3D moving targets via raw videos recorded by a set of cameras. The approach is composed of three key blocks: (1) a target tracking module that extracts plane pixel motions of the moving target in each video, (2) a Rodrigues' rotation formula-based coordinate transformation learning module that reconstructs the 3D coordinates with respect to a predefined reference point, and (3) a spline-enhanced library-based sparse regressor that uncovers the underlying governing law of dynamics. This framework is capable of effectively handling the challenges associated with measurement data, e.g., noise in the video, imprecise tracking of the target that causes data missing, etc. The efficacy of our method has been demonstrated through multiple sets of synthetic videos considering different nonlinear dynamics.
fields
cs.CE 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
MLLM-based Discovery of Intrinsic Coordinates and Governing Equations from High-Dimensional Data
VER is a multimodal-LLM framework that extracts low-dimensional coordinates from video and, together with SINDy, discovers governing ODEs without fine-tuning the LLM.