VLA-World improves autonomous driving by using action-guided future image generation followed by reflective reasoning over the imagined scene to refine trajectories.
Title resolution pending
10 Pith papers cite this work. Polarity classification is still indexing.
citation-role summary
citation-polarity summary
roles
method 1polarities
use method 1representative citing papers
KinemaForge jointly infers part geometry, joint topology, and parameters from RGB-D sequences using a kinematic graph and differentiable dynamics, then verifies with an energy residual loss, reporting lower joint errors and reduced simulation drift than PARIS and Ditto baselines.
Introduces MTRS task, MTRefSeg-21K benchmark of 21K image-text-mask triplets, and MTRefSeg-R1 LVLM baseline that outperforms standard models via two-stage change-aware training.
FSDrive uses a generated future scene frame as visual spatio-temporal CoT to improve VLA models for safer autonomous driving trajectory prediction.
OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.
PASE is a neuro-symbolic self-healing system that synthesizes LLM recovery plans, verifies them in simulation, and uses DRL to optimize prompts, claiming over 40% faster recovery on cloud fault data.
AtmoFuseNet fuses multi-view sky cameras, millimeter-wave radar, and ceilometer data via hierarchical cross-attention, variational refinement, and motion estimation to produce 4D cloud microphysical fields and wind with reported MAEs of 0.026 g m^{-3} LWC and 1.18 m s^{-1} wind speed.
Gradient boosting with conformal prediction and mutual-information stability selection yields NAFLD risk predictions with 91.3% empirical coverage at 90% nominal level and AUROC 0.91 on multicenter Chinese data.
Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.
VOTE-RAG applies retrieval voting across diverse queries and response voting across independent generations to mitigate hallucination-on-hallucination in RAG, matching or exceeding complex baselines on six benchmarks with a parallelizable design.
citing papers explorer
-
Learning Vision-Language-Action World Models for Autonomous Driving
VLA-World improves autonomous driving by using action-guided future image generation followed by reflective reasoning over the imagined scene to refine trajectories.
-
URDF Synthesis from RGB-D Sequences via Differentiable Joint Inference and Energy-Consistent Verification
KinemaForge jointly infers part geometry, joint topology, and parameters from RGB-D sequences using a kinematic graph and differentiable dynamics, then verifies with an energy residual loss, reporting lower joint errors and reduced simulation drift than PARIS and Ditto baselines.
-
An Open-Source Benchmark and Baseline for Multi-temporal Referring Segmentation
Introduces MTRS task, MTRefSeg-21K benchmark of 21K image-text-mask triplets, and MTRefSeg-R1 LVLM baseline that outperforms standard models via two-stage change-aware training.
-
FutureSightDrive: Thinking Visually with Spatio-Temporal CoT for Autonomous Driving
FSDrive uses a generated future scene frame as visual spatio-temporal CoT to improve VLA models for safer autonomous driving trajectory prediction.
-
Towards Realistic Open-Vocabulary Remote Sensing Segmentation: Benchmark and Baseline
OVRSISBenchV2 expands open-vocabulary remote-sensing segmentation evaluation to 170K images and 128 categories, and Pi-Seg uses positive-incentive noise to improve transfer on that harder benchmark.
-
Safe and Adaptive Cloud Healing: Verifying LLM-Generated Recovery Plans with a Neural-Symbolic World Model
PASE is a neuro-symbolic self-healing system that synthesizes LLM recovery plans, verifies them in simulation, and uses DRL to optimize prompts, claiming over 40% faster recovery on cloud fault data.
-
Cross-Modal Hierarchical Fusion for from Multi-Sensor Ground Observation
AtmoFuseNet fuses multi-view sky cameras, millimeter-wave radar, and ceilometer data via hierarchical cross-attention, variational refinement, and motion estimation to produce 4D cloud microphysical fields and wind with reported MAEs of 0.026 g m^{-3} LWC and 1.18 m s^{-1} wind speed.
-
Conformal Risk Prediction for Non-Alcoholic Fatty Liver Disease Using Gradient Boosting with Distribution-Free Coverages
Gradient boosting with conformal prediction and mutual-information stability selection yields NAFLD risk predictions with 91.3% empirical coverage at 90% nominal level and AUROC 0.91 on multicenter Chinese data.
-
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents
Audio language models are benchmarked on five semantic and paralinguistic reasoning tasks to reveal limitations in handling spoken audio evidence, accent variation, and domain shifts.
-
Mitigating Hallucination on Hallucination in RAG via Ensemble Voting
VOTE-RAG applies retrieval voting across diverse queries and response voting across independent generations to mitigate hallucination-on-hallucination in RAG, matching or exceeding complex baselines on six benchmarks with a parallelizable design.