Pith. sign in

REVIEW 9 cited by

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09024 v1 pith:JTGL7GUW submitted 2024-12-30 cs.CV cs.HCcs.RO

classification cs.CVcs.HCcs.RO
keywords robotreasoningnavigationperceptionsocialsocial-llavasociallyactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most existing social robot navigation techniques either leverage hand-crafted rules or human demonstrations to connect robot perception to socially compliant actions. However, there remains a significant gap in effectively translating perception into socially compliant actions, much like how human reasoning naturally occurs in dynamic environments. Considering the recent success of Vision-Language Models (VLMs), we propose using language to bridge the gap in human-like reasoning between perception and socially aware robot actions. We create a vision-language dataset, Social robot Navigation via Explainable Interactions (SNEI), featuring 40K human-annotated Visual Question Answers (VQAs) based on 2K human-robot social interactions in unstructured, crowded public spaces, spanning perception, prediction, chain-of-thought reasoning, action, and explanation. We fine-tune a VLM, Social-LLaVA, using SNEI to demonstrate the practical application of our dataset. Social-LLaVA outperforms state-of-the-art models like GPT-4V and Gemini, based on the average of fifteen different human-judge scores across 50 VQA. Deployed onboard a mobile robot, Social-LLaVA enables human-like reasoning, marking a promising step toward socially compliant robot navigation in dynamic public spaces through language reasoning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. HCSG: Human-Centric Semantic-Geometric Reasoning for Vision-Language Navigation

    cs.RO 2026-05 unverdicted novelty 7.0 of 10

    HCSG combines geometric forecasting of human pose and trajectory with VLM-generated semantic descriptions of intentions, fused into a topological map with a social distance loss, yielding 14% higher success rate and 3...

  2. Logic-Guided Socially-aware Robot Navigation World Model

    cs.RO 2025-10 conditional novelty 6.0 of 10

    NaviWM couples a spatial-temporal world model with a deductive chain-of-thought, formalizing social navigation rules as first-order logic, and reports improved success and lower violation rates in simulated crowded na...

  3. HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

    cs.RO 2025-08 reject novelty 6.0 of 10

    HALO learns a vision-based navigation reward from human preference rankings on egocentric video, and an IQL policy using it beats several baselines in 10-trial real-world tests.

  4. Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Conditional VLM reasoning triggered by personal-space violations improves social-navigation success by up to 20 points over RL-only baselines while keeping most steps on a fast RL policy.

  5. Learning Robot Visual Navigation in Crowds via Intention-Aware Scene Representations

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    iCrowdNav encodes egocentric visual observations with occupancy features and human pose intentions to improve DRL policies for crowd navigation, showing better performance than baselines in experiments and real-world tests.

  6. Act on What You See: Unlocking Safe Social Navigation in Vision-Language-Action Models

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    SALSA aligns social features and adds future-risk signals in VLA models to cut near-collisions by 86.4% and raise social accuracy from 53% to 93% on SCAND and real robots.

  7. MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments

    cs.CV 2025-12 reject novelty 5.0 of 10

    MUSON is a curated dataset of egocentric navigation images with chain-of-thought labels on which the main text reports Qwen2.5-VL-3B reaching 0.8625 action accuracy, while the arXiv abstract reports a different 10,110...

  8. AutoSpatial: Visual-Language Reasoning for Social Robot Navigation through Efficient Spatial Reasoning Learning

    cs.RO 2025-03 unverdicted novelty 5.0 of 10

    AutoSpatial improves VLM spatial reasoning for social navigation by combining minimal manual supervision with auto-labeled VQA pairs and hierarchical training, showing gains up to 20.5% in action prediction over baselines.

  9. Trust Through Transparency: Explainable Social Navigation for Autonomous Mobile Robots via Vision-Language Models

    cs.RO 2025-04 unverdicted novelty 4.0 of 10

    Multimodal explainability module using vision-language models and heat maps enables robots to generate natural-language summaries of navigation observations, with n=30 user studies showing majority preference for real...

Pith tools