Pith. sign in

REVIEW 1 cited by

Semantic Map-based Generation of Navigation Instructions

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.19603 v1 pith:DUYRAGTF submitted 2024-03-28 cs.CL cs.AIcs.CV

classification cs.CLcs.AIcs.CV
keywords generationnavigationsemanticinstructionsmapsimagesinstructionpanorama
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We are interested in the generation of navigation instructions, either in their own right or as training material for robotic navigation task. In this paper, we propose a new approach to navigation instruction generation by framing the problem as an image captioning task using semantic maps as visual input. Conventional approaches employ a sequence of panorama images to generate navigation instructions. Semantic maps abstract away from visual details and fuse the information in multiple panorama images into a single top-down representation, thereby reducing computational complexity to process the input. We present a benchmark dataset for instruction generation using semantic maps, propose an initial model and ask human subjects to manually assess the quality of generated instructions. Our initial investigations show promise in using semantic maps for instruction generation instead of a sequence of panorama images, but there is vast scope for improvement. We release the code for data preparation and model training at https://github.com/chengzu-li/VLGen.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. NAVCON: A Cognitively Inspired and Linguistically Grounded Corpus for Vision and Language Navigation

    cs.CL 2024-12 conditional novelty 6.0 of 10

    A new corpus adds 236,316 navigation concept annotations and 2.7 million aligned video frames to the R2R and RxR vision-language navigation datasets.

Pith tools