Perfect discrimination can still make certified admission impossible
Telling two states apart and certifying an action need different sensing designs — and different audits.
Robotics
Roughly includes material in ACM Subject Class I.2.9.
sort pith recommended most recent
Telling two states apart and certifying an action need different sensing designs — and different audits.
Directly predicts 29-DoF commands from RGB and text, achieving zero-shot sim-to-real transfer on a Unitree G1 robot.
DeCAL unifies understanding, imagination, and action generation with contact-aware tactile fusion to achieve state-of-the-art dexterous…
Reduces crashes by over 99% in simulation and runs at 60 Hz on a real racecar.
· “Online, Reachability-Aware, Sampling-Based Motion Planning”
Controlled experiments show ground-truth occupancy speeds coverage but not final coverage; a filter recovers targeted failures without…
· “Rethinking Learned Occupancy in Autonomous Active Mapping with Observation-Gated Filtering”
Field tests show a simple particle-spreading heuristic reduces peak error after communication dropouts without harming nominal performance.
· “A Distributed Consensus Particle Filter for Target Tracking using Autonomous Surface Vessels”
DYAD records 851 co-located assistance events linking requests, verbal and physical help, triggers, and outcomes during gearbox assembly.
· “DYAD: A Multimodal Dataset of Co-Located Human Assistance”
A new dataset and two geometry-aware modules bridge the angular-to-Cartesian gap for autonomous driving perception.
· “Spheriverse: 3D Scene Understanding from Spherical Observations in the Wild”
PlannerForge outperforms specialized tools on generation, selection, modification, and enhancement without fine-tuning.
· “PlannerForge: LLM Agents for Scenario-Based Testing of Motion Planners in Autonomous Driving”
A learned dynamics model augmented with differentiable contact detection enables a hybrid MPPI controller to navigate ramps, narrow…
· “Model Predictive Control of Tensegrity Robots via Contact-Aware Graph Neural Dynamics Model”
New method encodes arbitrary bits as subtle noise in robot behavior, allowing auditors with a shared key to recover intent from remote…
New visible-reachable metric cuts dual-target grasping time by 17% and energy by 19%.
· “Visible-Reachable Workspace for Perception-Aware Humanoid Design”
FRAME decomposes language queries and learns per-attribute probes, enabling fast and accurate multi-attribute object recall from fixed…
· “FRAME: Factored Retrieval via Attribute Readouts for Object-Centric Scene Memory”
CAST uses alternating planner and learned-policy transitions to train a state-value critic, boosting sample efficiency on 14…
A feed-forward framework parses objects and generates complete textured meshes without iterative optimization or manual prompts.
· “FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute”
Chamber-by-chamber perturbation and an autoencoder let untethered soft robots detect damage without extra hardware.
· “Real-time Puncture Detection and Recovery for Pneumatic Soft Actuators”
A safety-decoupled framework guarantees collision-free cooperative navigation under time-varying topologies.
· “Graph-Based Safe Reinforcement Learning for Multi-Agent Systems with Time-Varying Topology”
Keeps gradients reliable at 0.1 s steps: 8,192 worlds on one GPU, control tasks MJX cannot solve.
· “Ostrich: Taking Large Strides Through Stiff Contact in Differentiable Dynamics”
Same instrument, same eye model, same targets: swapping in a robot changed every movement measure.
FOCI Policy extracts short interaction segments and models object-relative trajectories, achieving strong data efficiency and…
· “FOCI Policy: Focus on Object-Centric Interactions for Relational Manipulation Policies”
Replacing sensor range with distance to the robot body in LiDAR features yields a 28-point gain, though seed selection limits the result.
· “DCLP++: Learning to Navigate with Footprint Clearance and Relative Motion”
Across three offline datasets, Bayesian neural nets outperform deterministic variants; an online study shows hierarchical explanations…
With just 10% labeled contacts on the new sensor, no retraining or paired calibration is needed.
· “BIFTA: Brain-Inspired Few-Shot Tactile Adaptation for Unknown Sensors”
A novel extension of soft actor-critic handles mixed actions and graph states, achieving real-robot closed-loop assembly.
· “Learning to build covering structures with continuous adjustments”
A product of successor value and pseudocount jointly scores novelty and reachability, outperforming tuned additive methods with fewer…
Occupancy-weighted stage descriptions beat current-step labels on LIBERO, RoboTwin, and MolmoSpaces; one variant declines.
· “CASD: Chunk-Aligned Semantic Distillation for Multi-StageRobot Manipulation”
Omnidirectional MFVINS with learned depth averaged 0.21 m error vs 1.39 m for monocular VINS and 2.61 m for MCVIO.
· “MFVINS: Multiple Fisheye Camera-Based Visual Inertial System”
New method gives robots a statistical way to know when their terrain understanding is unreliable, enabling safer autonomous navigation in…
Coarse region wedges set the route, fine viewpoints fill the gaps, and rough terrain gets mapped in about half the time.
· “TASG-Explore: Traversability-Aware Sector-Guided Exploration for Ground Robot on Uneven Terrain”
PGMT uses motion-conditioned terrain glimpses to adjust footholds and posture, enabling a Unitree G1 to handle stairs, rough ground, and…
· “PGMT: Perceptive General Motion Tracking for Humanoid Robots”
A robot hand that picks its next view from uncertainty beats fixed spin schedules on a 30-second budget.
· “AURORA: Active Uncertainty-Driven Re-Orientation for In-Hand Reconstruction”
Tap-Pull and Micro-Pulse, tuned by Bayesian optimization, turn passive elasticity into sustained, uniform drum sounds.
· “Multi-bounce Drum Roll with Optimized Active Tricks to Leverage Soft Embodiment”
By storing object relationships even after they leave view, SafeMem lifts safe success rate from 0.21 to 0.59.
· “Safe Task Planning with Long-Term Graph Memory for Embodied Agents”
AirAnchor bridges local and global space via shared object anchors, outperforming all zero-shot baselines on AerialVLN-S.
FocusPool learns task-progressive masks from intermediate layers, needing half the training data of baselines.
· “Localized Visual Feature Aggregation via Focus Pooling for Visuomotor Policies”
By selecting one IK solution per cell, the method reduces joint movement and avoids revisiting sub-cells.
· “Coverage Path Planning for Redundant Manipulators using Generalized Spanning Trees”
Enforcing verticality and coplanarity on RGB wireframes achieves 88% sequential success at 0.1 m on Gibson, outperforming depth‑based F3Loc.
· “GALoc: Gravity Aligned Wireframes for Depth-Free Monocular Floorplan Localization”
Automatically generates millions of bimanual manipulation demonstrations from user images, improving real-world transfer.
· “RoboCousin: Build Your Own Simulation Playground for Robust Bimanual Robotic Manipulation”
Real-world tests show camera-LiDAR-radar combinaton beats any pair in speed, range, and accuracy, enabling safe high-speed overtaking.
· “A Multi-Modal Perception Pipeline for Object Detection and Tracking in Autonomous Racing”
New benchmark shows that reusing stale scene representations cuts success rate by up to half in changing environments.
· “EvoNav-Bench: Benchmarking Lifelong Navigation in Evolving Environments”
Attackers replace telemetry before encryption; a Franka arm was hijacked 100% of the time while verifiers saw normal data.
· “Seeing is Not Believing: Breaking the Physical-to-Digital Trust Boundary in Robotics”
CALIPER test shows varying camera and clutter is needed to separate models that truly infer mass and friction from those that don't.
· “CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations”
Predicting consistent 2D points across cameras and triangulating them gives robots explicit 3D motion cues for generalisation.
· “3DWay: Generalizing Robot Manipulation via 3D Consistent Waypoints”
A closed-loop framework with simulation and critique achieves 96.2% validity and produces three working grippers.
· “Bridging Language and Physics: Automated Design of Continuum Robots with Large Language Models”
Short-term scene memory and long-term experience combine to beat larger models on driving benchmarks.
Fiber-optic clip captures contact force and texture vibrations for robot teaching, wet or dry.
· “TacClip: a clip-on sensor measures dynamic contact forces without covering the fingerpads”
A dual-layer spatial memory turns flickering AI sightings into steady search guidance, with best marks on the UAV-ON benchmark.
· “Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation”
OmniNav keeps memory, target belief, and reachability in sync, lifting real pick-and-place success from 53% to 72%.
· “OmniNav: Robust Long-Horizon Target Navigation in Dynamic Environments”
A tri-action policy raises dashcam AUROC from 0.610 to 0.689 and catches 89% of takeovers within five seconds.
· “Observe Before You Alert: Adaptive Driver Alerting with Vision-Language Models”
Grouping failed episodes into recurring modes and targeting each new demonstration yields the largest edge at the smallest budget.
· “DISEIL: Demonstration Distillation for Sample-Efficient Imitation Learning”
mjorbit couples MuJoCo with Encke's method in an orbit-following frame, matching Basilisk at 1e-7 m while running 1000× faster.
Decoupling the operator's gaze from robot motion speeds search, recovery, and bimanual tasks in VR teleoperation.
· “SPOT: Spatial Perception-Oriented Long-Horizon Humanoid Teleoperation”
Leveraging planned motion trajectories as exogenous variables, it outperforms state-of-the-art on real and simulated aperiodic benchmarks
· “A Multimodal Label Forecasting Method for Aperiodic Visuo-Motor Time Series”
M3-Tele combines force and tactile feedback to produce cleaner, learnable demonstrations for mobile manipulation policy learning.
Real-time semantic map drives vibra feedback; users align early without time penalty
CLIP picks the few latent cells that matter; a tuned diffusion model fills the rest, even on images never seen before.
· “Foundation Models for Generalizable Semantic and Goal-Oriented Communication”
The instruction knows what matters: sending 32 of 512 image tokens costs just 1.5 points of task success.
· “ComVLA: Communication-Aware Split Inference for VLA Models in 6G-Connected Robotics”
DEX-X learns visual-tactile manipulation from video and transfers zero-shot to a real hand-arm system, achieving 93% cube-picking success.
· “Dex-X: Learning Visual-Tactile Dexterous Manipulation From Human Videos with Simulated Interaction”
Predictive safety layer and a deadlock-breaking protocol allow pre-trained multi-agent policies to adapt safely to unseen obstacles.
· “Decentralized Safe Multi-Agent Reinforcement Learning via Predictive Shielding”
Retrieved demonstration snippets give a frozen text-only policy a 19-point edge on RoboTwin 2.0.
A 0.15 m four-step ceiling walk shows that commanding thrust derivatives instead of thrust kills the spikes that break contact.
PACE encoder learns domain-invariant features by predicting joint transitions, enabling direct policy deployment without real data.
· “Zero-Shot Sim-to-Real Contact-Rich Assembly via Proprioception-Anchored Cross-Modal Pretraining”
PhysReal's hybrid material model predicts how toys and cloth deform under new pushes better than four baselines.
· “PhysReal: Learning Real-World Deformable Object Physics via Hybrid Constitutive Modeling”
Fixing hole radius and rectangle layout eliminates need for dense LiDAR, extending usable standoff range 1.5×.
· “P²Calib: Utilizing Pattern Priors for LiDAR-Camera Extrinsic Calibration”
First demonstration that learning-based methods can plausibly extend HD maps beyond sensor range on simple roads, offering a path toward…
Chunked continuous actions from one demonstration rival per-task fine-tuning in simulation and on real bimanual arms.
· “ContextFlow: In-Context Flow Matching for Robot Manipulation”
Diagnosis over 25 starts finds occupancy fixes mainly speed exploration; a new filter rescues two hard cases.
· “Diagnosing and Dynamically Filtering Occupancy World Models for Active Mapping”
Simulation shows a win-win when traffic coverage is added to route planning with the right weighting.
· “RoboSense: Leveraging Robotaxi Fleets as Drive-by Sensors for Urban Traffic Monitoring”
A hybrid neural representation turns discrete OCT into a continuous tissue field, speeding robotic path planning and sparse reconstruction.
· “OCTN: Neural OCT Representations for Robot-Guided Precision Intervention”
A simplified model and off-the-shelf parts let a four-legged robot turn and tilt without propellers.
· “Design and Attitude Control of an Underwater Quadruped Robot”
SkillX unifies three soccer skills in a single neural network and transfers from simulation to a real robot.
· “SkillX: Unified Multi-Skill Policy Learning for Humanoid Soccer”
Augmenting VLM outputs with OpenStreetMap and sensor data yields building F1=0.83 in outdoor tests.
· “Physico-Geospatial Grounded Scene Interpretation for Mobile Robotics”
Diffusion planner with a 512-dim recurrent state covers unseen spaces in simulation and flies an Elios 3.
· “CAVEAT: Recurrent Multimodal Diffusion Planning for Mapless Aerial Exploration”
Outperforms vision baseline by 8.67 points under combined position and viewpoint shifts.
A context-conditioned discrete skill prior enables reusable interaction skills without task-specific reward engineering, achieving 88%…
· “Unifying Physics-Based Humanoid Interaction with a Context-Conditioned Interaction Prior”
Walls and floors stop being re-fused every frame; mesh error stays under 0.19 cm while TSDF cost drops.
Simulated in 500 episodes, transferred directly to a real robot with 100% success over 20 trials.
A control proof and simulation show a camera drone can replace a costly fixed-wing plane for training and testing.
· “Dynamic System Emulation: Fixed Wing Dynamics on a Multicopter”
A learned verifier uses past images, actions, and states to detect failures and issue stage-aware fix prompts, boosting success without…
A 50-person study finds no reliable difference in perceived agency between human-operated and AI-generated robot behavior.