Pith. sign in

Paper Citation Record · LEDGER

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2506.08566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08566 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:12:12.349149Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T23:02:14.913189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T23:02:52.667737Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact10
  • verified fuzzy21
  • unresolved25
  • parse uncertain0
  • malformed identifier7
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b4c694c-d47f-4296-a2a7-f1778dedafb3 · outbound

This paper cites Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.021740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.021740Z digest=sha256:0fbe1353848a936aa5e0fed5e3228741c99c51b45cbc738a0cac3a5a75e5538d

Observation 96807f12-72d1-4b10-ae32-2c537c85d637 · outbound

This paper cites Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.027146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.027146Z digest=sha256:5a368769dcd5f268a1c32d336f24cbc0ab5db2361e4f15ac5ec267e72d3f2420

Observation eacef57d-aabb-412c-a4c9-80b3bdc53d23 · outbound

This paper cites Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.032305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.032305Z digest=sha256:70a74a3ea5fed3e92f57afb047a6b069cb917dc3be248969bda616c32984b122

Observation a2ccec0e-972c-4408-81c0-c0f9a75152c0 · outbound

This paper cites NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P

Reference 4

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:14.075459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.036976Z digest=sha256:c0b30811264d890a956bdc0e6e34bb12ea7fd165f13c5cd177e9675acdb91386

Observation 34cf6f45-2840-4dc4-8083-16501565f751 · outbound

This paper cites Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.654668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.041272Z digest=sha256:3c3a1a49a3b3ab6891a20774bd489e80b2ff49188719de9fcc4932c14469559f

Observation 9ac936c7-125b-458d-b108-78302e980c76 · outbound

This paper cites TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.050547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.050547Z digest=sha256:69658f6bd62cfeb771f10ce93d83d17711f78376e40f81cd174fff15d46c8137

Observation 4c521f32-b680-4250-a58c-cdfd0af18af6 · outbound

This paper cites Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.641035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.055232Z digest=sha256:be3306a842396b28496193322824ef48fb6b53a0306a1d3a6fc55dcdee61c842

Observation 72778319-bebe-4dcc-9303-e3e108bda61f · outbound

This paper cites Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.059807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.059807Z digest=sha256:a3e066d9859f101d3373658cbf5a2978bc9537801530d3682c511f1c78bdb02b

Observation 4800851c-3b88-4622-9bab-b5fd9a8a98b2 · outbound

This paper cites Unifying vision-and- language tasks via text generation, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unifying vision-and- language tasks via text generation, in: ICML, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.627226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.068175Z digest=sha256:7d9c228a34395ce35239130acda653121b5af6a9d9faae5a972d29679f4587c2

Observation c0100153-0501-4fb7-a125-7fce788e1b16 · outbound

This paper cites Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.072754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.072754Z digest=sha256:7b7f4142ea926e60c1288fa6296a9c26bd10dca050f3a78923bd2c4f802c94f2

Observation 88915f0b-fdae-43db-8e09-6efb5b734a8c · outbound

This paper cites Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.613994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.081842Z digest=sha256:a4075e60ac1680e1ed0d18aa159b6019b0b454e74f6886ddbe5144aea4ecdeb7

Observation e6f20fcc-cd24-44e8-b6f8-e65156778454 · outbound

This paper cites Visionformobilerobotnavigation: A survey.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Visionformobilerobotnavigation: A survey

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.517455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.090129Z digest=sha256:da69a347eb1e44c5fb412cfe1cc5edf67fbe1675c4a3a28863bc64f20fdf46d3

Observation ea206fc4-b2f1-43d0-9691-7a816cbfdb65 · outbound

This paper cites Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.094146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.094146Z digest=sha256:caa59cd7d3f3be91222d4f7f2016dc885551f717e1734a3985a1cc32b04c1dbf

Observation 8da7b90f-f272-4703-9169-ab8ce3500280 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.495451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.098369Z digest=sha256:7482b90e21b215038a2f16c66f089c9f242bd7416520a0a632b037c06a6d5364

Observation 3689e419-a13b-4ab2-b2af-a6f0384b1ffc · outbound

This paper cites Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.599570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.102791Z digest=sha256:adface74f0c84122aec178f8bffbcdb82e55ae86ba5465ecd38714998545f4b4

Observation 025a00c6-43d4-4456-aa90-5ecf5772357c · outbound

This paper cites Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.111576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.111576Z digest=sha256:8061860a349afb4e72f073ad08a621ccd3ffcd46b9b511fe34907c8de7de3ccc

Observation 752621a0-ee27-464f-9529-1c43697b88c4 · outbound

This paper cites Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.586568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.124012Z digest=sha256:b1bae370e72df7b937ecb7e1586d3dbd902205a5f266208d33725f770378fd9d

Observation 8d989065-f2ba-40ba-97dd-c59da456644d · outbound

This paper cites Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.572587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.128097Z digest=sha256:c94477979bc65f37ed9a0e4c548d62759f214c3ab52871fd8659d135e833ac93

Observation 966cfa01-4b5d-4b6f-98cb-97e5af460351 · outbound

This paper cites Sub-instruction aware vision-and-language navigation, in: EMNLP, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Sub-instruction aware vision-and-language navigation, in: EMNLP, pp

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.473819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.132307Z digest=sha256:f564f2fe0cc52a5bf757acf8b5d92c94241ebf067652f288924dd56fa78dbddd

Observation 5ea3fefe-2247-452c-847e-f97c09506236 · outbound

This paper cites VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.136388Z digest=sha256:171b13d05c4c6e32df0d6db1cc416e74ce2abdaade123301aed4d2258dcee305

Observation 76cec562-f291-462e-a3fd-11258133fb05 · outbound

This paper cites Transferable representation learning in vision-and- language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Transferable representation learning in vision-and- language navigation, in: ICCV, pp

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.140285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.140285Z digest=sha256:73db8bc0147515e858f1780a1522d2d6d85b1b0082b6c7d22bd84039b7fa6a64

Observation d90a169e-7aca-4fc8-a8c3-7a16064af85c · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.149291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.149291Z digest=sha256:daa1b63d59ff725c4176890455f41c8899f54baca07cf84fa57fad0146e0636c

Observation aaae132a-7902-4e99-9330-79976e2d8237 · outbound

This paper cites Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.559165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.153917Z digest=sha256:e52664b1852c8e590601188a7ca380e5e97cfa3ddc86a06eae36f0d89b75d21a

Observation 6057cf87-c913-48cb-8dda-9a4383e98b0a · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.262137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.161951Z digest=sha256:2b02b9046fc4cac061bd04726800dc6a8e09a3ce9e0eefb01eaf52687b2948ed

Observation f1893db0-14cf-41a9-a385-1fb30c8d9fcc · outbound

This paper cites Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.545422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.167048Z digest=sha256:707c0ee883ed4fafd2e1d33bdf1fabc2778d2089af23aab06b04c18b0ab547c5

Observation 7cbc4aa2-365c-4f2b-842e-70b1e23089ef · outbound

This paper cites Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.451747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.171799Z digest=sha256:1e24cc0555268b0c3b824f91429c3a305700afffea22722610928cd476b46b98

Observation 8be528bc-5df8-494e-98d0-835d505db6f9 · outbound

This paper cites Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.532183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.175858Z digest=sha256:c5ca13c8f6b357dfc0cf393d539596b852cca210301f7d2bc7649f441c68c638

Observation dac5984a-7651-4248-a27b-e5b54a09d0e3 · outbound

This paper cites KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.184569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.184569Z digest=sha256:3338c21f8aff8b09e4646bc5e457575eb3d2dad3aa82015b5feb31abac10a6bb

Observation 214a6609-ca0c-460b-8f6d-37c8d98f1a0e · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.518389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.188385Z digest=sha256:7a3195b3de5571a62c0252cab71b4342c95a5102fb6ffd560859ae4b9e3ac85b

Observation f485b723-0a93-459b-837d-0e5adc672e0f · outbound

This paper cites Learning vision-and-language navigation from youtube videos, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning vision-and-language navigation from youtube videos, in: CVPR, pp

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.192628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.192628Z digest=sha256:d96bac2b281973dbf6aba50c5eb188d28d574c5da4b83fbf5f413e679d04488d

Observation 8794a4ed-015e-4f2b-a699-0e8aee49458f · outbound

This paper cites Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.505709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.200454Z digest=sha256:b41afd11cb0f206921691d8ed32edfdbfba07794e514df286f9aefc567bca1e2

Observation 65b1f3e3-261b-4320-9e2d-0ed0cfd45906 · outbound

This paper cites Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.181081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.204524Z digest=sha256:34206abd062ed1e129af1a6f297879322801a61cd10d6c73fcaf0dcfa4a64af8

Observation 0187b010-9da3-4203-82d9-cbc16ca6e155 · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.492370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.208375Z digest=sha256:c513f16bce82153c5ce4be578683a95ff9c0cc21ede7c928f9f677a58fb88b30

Observation d50d4463-a269-43cd-9579-b13975f35f18 · outbound

This paper cites SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.478509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.212340Z digest=sha256:1a1294b997ba41faf9b29d211e2259751a6314935c2814f1943713f9630cb67a

Observation 9396e758-f39a-412f-bd9a-ccae3a5eaafa · outbound

This paper cites Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.220240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.220240Z digest=sha256:cbfc0640bf7c4d1c790e69e724525dbe47a4e9c03ed9510bc12140efa47d4c23

Observation b2052fa8-2f1e-4b64-8f95-d40c067ac602 · outbound

This paper cites Teach: Task-driven embodied agents that chat, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Teach: Task-driven embodied agents that chat, in: AAAI, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.450898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.224781Z digest=sha256:f8b21bb465bec00a5bc757a385fd3db12d6eb1506fcf6bc5f8d1947c02ea15db

Observation 88509fdc-81c0-4ebc-bac5-551fa243ac51 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.235094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.235094Z digest=sha256:217c615485760f2d0224391b27e99f35627a7aaaa8fa96151ef9ec5316c47324

Observation 661964f5-77b7-4b2e-9318-550233c04d4a · outbound

This paper cites The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.436917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.239091Z digest=sha256:a07de4b59cada575b8adb7cc485a660867a9fc57d56782dc593a56f2c23c08a6

Observation 23b9ec66-062c-4738-82ce-742964fb1c6c · outbound

This paper cites Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.419946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.247499Z digest=sha256:ac4db0151b959ea2764a505bb4e314cb97cd1899342429aaae609cb35e33719a

Observation 2c28b650-edb2-45fc-bfd5-f82a69f624fe · outbound

This paper cites HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.256573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.256573Z digest=sha256:9faf890826362e1e25feaf3f43e6532d11d5ea1eae582f64f085a5a79756cccf

Observation 615a9ec6-fb64-4bf0-b055-d42276909213 · outbound

This paper cites Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.422608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.260608Z digest=sha256:e83c0c145c25a74d04632d75f120927e1a7a2d56f6c038cd8e39905d9db2f6df

Observation 0fea42e3-ee92-4f61-9f2a-ff3cf639a98c · outbound

This paper cites Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.407526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.268752Z digest=sha256:8baa359bdae79349f486aeecf14d44ab23e9c34b3513fd7b942c8b2d23a771fe

Observation bb77a352-9934-4565-ae53-20ebda4ace23 · outbound

This paper cites Contrastive search is what you need for neuraltextgeneration.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Contrastive search is what you need for neuraltextgeneration

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.393190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.272856Z digest=sha256:c0883664947c19c0c4d43548ecd9171e2aa1aa9ab249d133ef19eaea8677c99b

Observation aecb1bbf-1e92-4d93-933c-7851a130c470 · outbound

This paper cites Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp

Reference 55

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.405485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.277578Z digest=sha256:577a76367a02015e70f75c0d963ac80502fabba30bd5c74e56d64cb761918300

Observation 813c2439-0122-4ab6-935c-edc2581e9a80 · outbound

This paper cites Vision-and-dialog navigation, in: CoRL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-and-dialog navigation, in: CoRL, pp

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.281829Z digest=sha256:2dd1eed62901dd44b7d555e20440605f582bbed93a4a02bf7568374045ea78b7

Observation fa53d099-7623-4a10-9e7a-de62d4e3d867 · outbound

This paper cites Cider: Consensus- based image description evaluation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Cider: Consensus- based image description evaluation, in: CVPR, pp

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.285974Z digest=sha256:5fed213e40e83137d0d3895c7261ab44d442a692088cd4b74d1188fc17f6b50c

Observation 04a3101a-bf63-467a-b49b-d5e2dec4bf35 · outbound

This paper cites Show and tell: A neural image caption generator, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show and tell: A neural image caption generator, in: CVPR, pp

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.290006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.290006Z digest=sha256:6fe6c9e1ff91c1026f12df57bf985566d09d16657b29f4a90e5cba887a6ffa9b

Observation 3ab685b3-0fb2-443e-b9da-57a4ddd0134d · outbound

This paper cites Soft expert reward learning for vision-and-language navigation, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Soft expert reward learning for vision-and-language navigation, in: ECCV, pp

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.363551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.298071Z digest=sha256:c38a0d7d7b95cd0a27251005fdbf4675a34ca2585ac2019166ddadd5e3271e08

Observation 225f25e8-056b-4056-9d25-f6a7a05c2bf7 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations GIT: A generative image-to-text transformer for vision and language

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.347964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.303260Z digest=sha256:84eafb0fc4cca20fd5c6e0515dafd26a5312d1b0c328f6da36460a10def7b662

Observation 341561be-031f-4fea-af60-01fe927b83a4 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.320063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.311490Z digest=sha256:d9826e7a7f94305703d5e2c81fab1c4722da98eb3249c12f4217133f72483b8e

Observation 63c57a3a-3ba8-4d14-8b8f-16d07ccc25f9 · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.306372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.319606Z digest=sha256:b2e8748bba83156508e7c8e14b5c04afafdc444d7aadc66662d227085360c7aa

Observation 0609499b-c5eb-4540-b01f-9e7fb8b15770 · outbound

This paper cites Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.323819Z digest=sha256:38ac6324dbca09781c2398cc524dfb7e5434d3aa413d834fe935ce5064e4e25f

Observation 9edd9f04-e220-4576-822d-f24e52c0f593 · outbound

This paper cites Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.756522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.327661Z digest=sha256:86a308959a53085e6bfc638ed8e39ecacf6548aa896098dd82a7cfe483924f03

Observation 34bea04b-a9f1-4154-b860-46b8db0f2882 · outbound

This paper cites LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE

Reference 66

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.677546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.331590Z digest=sha256:86510155732b2c75223057bd8200586c459cb0a3f76c25432def2aee1d00ad9c

Observation 3ac4d94b-438e-41ff-acfd-bbfae8db2b06 · outbound

This paper cites Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.292112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.336239Z digest=sha256:a827f62be0490f37f1505477772a67cd3942a310857f38a5c313f31ef5ae41ed

Observation d4d4cf0d-a8ac-4376-9cf1-1a02599e6059 · outbound

This paper cites Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.340247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.340247Z digest=sha256:a1b90353d108a957d1b06c0319fe77597906385835a7d00d0ebc8836b61f6df8

Observation 6babf70e-4066-4e21-8f4c-c977d6cd47cd · outbound

This paper cites Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.345285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.345285Z digest=sha256:47a87ddfb85c4f767590d2fa89e7a68661325464ff6e6b65c5cd1f29ca2ece58

Observation ca6b9874-9a29-4ed5-86d9-409884f01816 · outbound

This paper cites Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.349149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.349149Z digest=sha256:ba59002977c4ab294467e3233bc3b19180ee6cea8f9d2e43a158b737081a4a1e

Observation 026645a0-97a9-4a3b-9725-f08d96dcf612 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 380

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.085959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.085959Z digest=sha256:8a203003e8246bb18328da64439e73b496cdf2e7cc19d079801e0a7e8eace595

Observation b910639c-c43c-4e4b-b389-5cbba8e5a53d · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 676

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.045804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.045804Z digest=sha256:7c066d6d8916f69b18da4461972e7629aeeda0f28c804baeeabfcf78144d1377

Observation cf621fab-ba65-4bc6-a139-d8eb16438d01 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 1644

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.243135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.243135Z digest=sha256:ed1baf2f48a80beddab465e73cd42fe9a7495b1db67147cd2f0dd6e3b67d21d7

Observation 538109e6-4675-4359-a24c-0044a6b7a330 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.333845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.307332Z digest=sha256:56ec4eccab1dc36e3dccf613267039defe5bf96c2962364c1e3effa1d7723013

Observation 160a8a3a-27e1-40c4-bc54-ab34e3c7420b · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2024

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:12.830343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.315636Z digest=sha256:325770bdebf8edde546cbfd2c653079d3cec59b3d878f5d491d5b2ccb04a1815

Observation 56374228-da84-4be0-885b-62f890fe8cef · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.229720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.229720Z digest=sha256:e69bfcdba50cdf7d0d2886968dd3c36a0826998e67893c2dbbd3c245926af53f

Observation 8fd95b62-9716-4c42-8769-55257bce25fb · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 7367

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.465023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:12:12.216264Z digest=sha256:d5e43e74e9a5f1090533fa8be695eda2b560b7d05425be4f1ac638d8b89ab469

Pith citing papers

Observation 134c4ef6-5447-4a21-bc5c-649dbd55a063 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:02:52.670999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:4f53c0833d3c9089eff56c2124a783da60ce522274cf91ec8e577bd08742fd1a