Pith. sign in

REVIEW 5 cited by

Social-LLaVA: Enhancing Robot Navigation through Human-Language Reasoning in Social Spaces

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09024 v1 pith:JTGL7GUW submitted 2024-12-30 cs.CV cs.HCcs.RO

classification cs.CVcs.HCcs.RO
keywords robotreasoningnavigationperceptionsocialsocial-llavasociallyactions
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Most existing social robot navigation techniques either leverage hand-crafted rules or human demonstrations to connect robot perception to socially compliant actions. However, there remains a significant gap in effectively translating perception into socially compliant actions, much like how human reasoning naturally occurs in dynamic environments. Considering the recent success of Vision-Language Models (VLMs), we propose using language to bridge the gap in human-like reasoning between perception and socially aware robot actions. We create a vision-language dataset, Social robot Navigation via Explainable Interactions (SNEI), featuring 40K human-annotated Visual Question Answers (VQAs) based on 2K human-robot social interactions in unstructured, crowded public spaces, spanning perception, prediction, chain-of-thought reasoning, action, and explanation. We fine-tune a VLM, Social-LLaVA, using SNEI to demonstrate the practical application of our dataset. Social-LLaVA outperforms state-of-the-art models like GPT-4V and Gemini, based on the average of fifteen different human-judge scores across 50 VQA. Deployed onboard a mobile robot, Social-LLaVA enables human-like reasoning, marking a promising step toward socially compliant robot navigation in dynamic public spaces through language reasoning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Logic-Guided Socially-aware Robot Navigation World Model

    cs.RO 2025-10 conditional novelty 6.0 of 10

    NaviWM couples a spatial-temporal world model with a deductive chain-of-thought, formalizing social navigation rules as first-order logic, and reports improved success and lower violation rates in simulated crowded na...

  2. HALO: Human Preference Aligned Offline Reward Learning for Robot Navigation

    cs.RO 2025-08 reject novelty 6.0 of 10

    HALO learns a vision-based navigation reward from human preference rankings on egocentric video, and an IQL policy using it beats several baselines in 10-trial real-world tests.

  3. Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Narrate2Nav uses Barlow Twins alignment to distill language-based reasoning from a large teacher into a small RGB-only navigation model, reporting lower trajectory error and higher goal-reaching success than four baselines.

  4. Think When It Matters: Conditional VLM Reasoning for Social Navigation with RL Policies

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Conditional VLM reasoning triggered by personal-space violations improves social-navigation success by up to 20 points over RL-only baselines while keeping most steps on a fast RL policy.

  5. MUSON: A Reasoning-oriented Multimodal Dataset for Socially Compliant Navigation in Urban Environments

    cs.CV 2025-12 reject novelty 5.0 of 10

    MUSON is a curated dataset of egocentric navigation images with chain-of-thought labels on which the main text reports Qwen2.5-VL-3B reaching 0.8625 action accuracy, while the arXiv abstract reports a different 10,110...

Pith tools