Pith. sign in

Paper Citation Record · LEDGER

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

As of 7 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2506.08566.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08566 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:12:12.349149Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T23:02:14.913189Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T23:02:52.667737Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact10
  • verified fuzzy21
  • unresolved25
  • parse uncertain0
  • malformed identifier7
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b4c694c-d47f-4296-a2a7-f1778dedafb3 · outbound

This paper cites Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Spatio- temporal dynamics and semantic attribute enriched visual encoding forvideocaptioning,in:CVPR,pp.12487–12496

Reference 1

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.021740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.021740Z digest=sha256:0fbe1353848a936aa5e0fed5e3228741c99c51b45cbc738a0cac3a5a75e5538d

Observation 96807f12-72d1-4b10-ae32-2c537c85d637 · outbound

This paper cites Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bottom-up and top-down attention for image captioningandvisualquestionanswering,in:CVPR,pp.6077–6086

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.027146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.027146Z digest=sha256:5a368769dcd5f268a1c32d336f24cbc0ab5db2361e4f15ac5ec267e72d3f2420

Observation eacef57d-aabb-412c-a4c9-80b3bdc53d23 · outbound

This paper cites Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: Interpreting visually-grounded navigation instructions in real environments, in: CVPR, pp

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.032305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.032305Z digest=sha256:70a74a3ea5fed3e92f57afb047a6b069cb917dc3be248969bda616c32984b122

Observation a2ccec0e-972c-4408-81c0-c0f9a75152c0 · outbound

This paper cites NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations NLTK: the natural language toolkit, in: Calzolari, N., Cardie, C., Isabelle, P

Reference 4

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:14.075459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.036976Z digest=sha256:e2235077e0f15772513689ea12468728628b44120c13cd44db1667d52731ad6c

Observation 34cf6f45-2840-4dc4-8083-16501565f751 · outbound

This paper cites Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Matterport3d: LearningfromRGB-Ddatainindoorenvironments,in:3DV,pp.667–

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.654668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.041272Z digest=sha256:fc92ef147b7a9ce2dd93d53c8d3929a5f9a7d434661c2b06949674486757e2e4

Observation 9ac936c7-125b-458d-b108-78302e980c76 · outbound

This paper cites TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations TOUCH- DOWN: natural language navigation and spatial reasoning in visual street environments, in: CVPR, pp

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.050547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.050547Z digest=sha256:69658f6bd62cfeb771f10ce93d83d17711f78376e40f81cd174fff15d46c8137

Observation 4c521f32-b680-4250-a58c-cdfd0af18af6 · outbound

This paper cites Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Historyawaremul- timodaltransformerforvision-and-languagenavigation,in:NeurIPS, pp

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.641035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.055232Z digest=sha256:3f91403a1bd8a4ff1e1756dd6a8f037ac1b59c4f8fd822fd8d4ffcb7a27c2079

Observation 72778319-bebe-4dcc-9303-e3e108bda61f · outbound

This paper cites Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning from unlabeled 3d environments for vision-and-language navigation,in:ECCV,pp.638–655

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.059807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.059807Z digest=sha256:a3e066d9859f101d3373658cbf5a2978bc9537801530d3682c511f1c78bdb02b

Observation 4800851c-3b88-4622-9bab-b5fd9a8a98b2 · outbound

This paper cites Unifying vision-and- language tasks via text generation, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unifying vision-and- language tasks via text generation, in: ICML, pp

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.627226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.068175Z digest=sha256:f046d33313db481b8619c888d94f4fbe39c91f61d926ab8d4df333a1ce2dcca1

Observation c0100153-0501-4fb7-a125-7fce788e1b16 · outbound

This paper cites Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Auxiliary fine-grained alignment constraints for vision-and-language navigation, in: ICME, pp

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.072754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.072754Z digest=sha256:7b7f4142ea926e60c1288fa6296a9c26bd10dca050f3a78923bd2c4f802c94f2

Observation 88915f0b-fdae-43db-8e09-6efb5b734a8c · outbound

This paper cites Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Meteor universal: Language specifictranslationevaluationforanytargetlanguage,in:Proceedings of the Ninth Workshop on Statistical Machine Translation, pp

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.613994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.081842Z digest=sha256:8dfe4faa53d54176b638a071f9410cb63eb24e71216e62d1b9dc576eb5daba98

Observation e6f20fcc-cd24-44e8-b6f8-e65156778454 · outbound

This paper cites Visionformobilerobotnavigation: A survey.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Visionformobilerobotnavigation: A survey

Reference 14

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.517455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.090129Z digest=sha256:eab55aec3a9755e544bd91497a39b7225f56c0d0c29875d5151c1472f45969fe

Observation ea206fc4-b2f1-43d0-9691-7a816cbfdb65 · outbound

This paper cites Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Aerial vision-and-dialog navigation, in: Findings of the Association forComputationalLinguistics,pp.3043–3061.doi:10.18653/V1/2023

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.094146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.094146Z digest=sha256:caa59cd7d3f3be91222d4f7f2016dc885551f717e1734a3985a1cc32b04c1dbf

Observation 8da7b90f-f272-4703-9169-ab8ce3500280 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 16

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.495451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.098369Z digest=sha256:80b55871b26584384035b9342f15b1d152815b8c32e9a7d902f161f3930399c5

Observation 3689e419-a13b-4ab2-b2af-a6f0384b1ffc · outbound

This paper cites Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Speaker-follower models for vision-and-language navigation, in: NeurIPS, pp

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.599570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.102791Z digest=sha256:f1c44961b209b5da5e57c6cef456a36f9d852971e94c4dac2fdb8953755efc95

Observation 025a00c6-43d4-4456-aa90-5ecf5772357c · outbound

This paper cites Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision- and-language navigation: A survey of tasks, methods, and future directions, in: ACL, pp

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.111576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.111576Z digest=sha256:8061860a349afb4e72f073ad08a621ccd3ffcd46b9b511fe34907c8de7de3ccc

Observation 752621a0-ee27-464f-9529-1c43697b88c4 · outbound

This paper cites Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Landmark-rxr: Solving vision-and-language naviga- tion with fine-grained alignment supervision, in: NeurIPS, pp

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.586568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.124012Z digest=sha256:70554149664fdfc56cc2dda6467b26b09666beee24709bd2d971794d2ffa8550

Observation 8d989065-f2ba-40ba-97dd-c59da456644d · outbound

This paper cites Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Lan- guage and visual entity relationship graph for agent navigation, in: NeurIPS

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.572587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.128097Z digest=sha256:9437784365891fcbe53aa88f0d723de94a1aa937a3cc4a74bdd86fd010dd2467

Observation 966cfa01-4b5d-4b6f-98cb-97e5af460351 · outbound

This paper cites Sub-instruction aware vision-and-language navigation, in: EMNLP, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Sub-instruction aware vision-and-language navigation, in: EMNLP, pp

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.473819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.132307Z digest=sha256:5b2ae8bb18f3d2f3b7a7c42f9b4cc030b8817058f6ceaf2818e317a0441841b7

Observation 5ea3fefe-2247-452c-847e-f97c09506236 · outbound

This paper cites VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations VLNBERT: Arecurrentvision-and-languageBERTfornavigation,in:CVPR,pp

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.136388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.136388Z digest=sha256:171b13d05c4c6e32df0d6db1cc416e74ce2abdaade123301aed4d2258dcee305

Observation 76cec562-f291-462e-a3fd-11258133fb05 · outbound

This paper cites Transferable representation learning in vision-and- language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Transferable representation learning in vision-and- language navigation, in: ICCV, pp

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.140285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.140285Z digest=sha256:73db8bc0147515e858f1780a1522d2d6d85b1b0082b6c7d22bd84039b7fa6a64

Observation d90a169e-7aca-4fc8-a8c3-7a16064af85c · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 28

Resolution
malformed identifier
no resolver link, observed 2026-08-07T05:12:12.149291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.149291Z digest=sha256:daa1b63d59ff725c4176890455f41c8899f54baca07cf84fa57fad0146e0636c

Observation aaae132a-7902-4e99-9330-79976e2d8237 · outbound

This paper cites Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Simple and effective synthesisofindoor3dscenes,in:AAAI,pp.1169–1178

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.559165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.153917Z digest=sha256:462784e2407293864f1f45e564337ec6bb57a960d886624b2d70b010ff5c43ee

Observation 6057cf87-c913-48cb-8dda-9a4383e98b0a · outbound

This paper cites Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Waypoint models for instruction-guided navigation in continuous environments, in: ICCV, pp

Reference 31

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.262137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.161951Z digest=sha256:b2f1d3140a32858cd45f8d2ea12800f66907565efc82caef86b712a98623f7d7

Observation f1893db0-14cf-41a9-a385-1fb30c8d9fcc · outbound

This paper cites Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Room- across-room:Multilingualvision-and-languagenavigationwithdense spatiotemporalgrounding,in:EMNLP,pp.4392–4412.doi:10.18653/ v1/2020.emnlp-main.356

Reference 32

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.545422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.167048Z digest=sha256:36c4fe231c57e3c03a96bf6625c100bd9eef85c12e24ef0a4ca5bcc161516c6b

Observation 7cbc4aa2-365c-4f2b-842e-70b1e23089ef · outbound

This paper cites Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unicoder- vl: A universal encoder for vision and language by cross-modal pre- training, in: AAAI, pp

Reference 33

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.451747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.171799Z digest=sha256:7786d74748a9297a1f8ab91e43b908d40bf1b3d03feb6072ab844c1e7e610462

Observation 8be528bc-5df8-494e-98d0-835d505db6f9 · outbound

This paper cites Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Panogen: Text-conditioned panoramic envi- ronment generation for vision-and-language navigation, in: NeurIPS

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.532183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.175858Z digest=sha256:bd45221a298256ef95a4f917952b26b7ac438ca923aa13e470064535176806f1

Observation dac5984a-7651-4248-a27b-e5b54a09d0e3 · outbound

This paper cites KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations KERM: knowledge enhanced reasoning for vision-and-language navigation, in: CVPR, pp

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.184569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.184569Z digest=sha256:3338c21f8aff8b09e4646bc5e457575eb3d2dad3aa82015b5feb31abac10a6bb

Observation 214a6609-ca0c-460b-8f6d-37c8d98f1a0e · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations ROUGE: A package for automatic evaluation of summaries,in:TextSummarizationBranchesOut,Barcelona,Spain

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.518389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.188385Z digest=sha256:98f1e7d9764bd9c17ad74bdee7d5fbc9758b7015fd77d959dde4f2b6fd7149c4

Observation f485b723-0a93-459b-837d-0e5adc672e0f · outbound

This paper cites Learning vision-and-language navigation from youtube videos, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning vision-and-language navigation from youtube videos, in: CVPR, pp

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.192628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.192628Z digest=sha256:d96bac2b281973dbf6aba50c5eb188d28d574c5da4b83fbf5f413e679d04488d

Observation 8794a4ed-015e-4f2b-a699-0e8aee49458f · outbound

This paper cites Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Self-monitoring navigation agent via auxiliary progress estimation, in: ICLR

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.505709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.200454Z digest=sha256:ea88854d346f27d45d13760d02e8a8c037fdd009c1c0dfa978b60578a4f61fcb

Observation 65b1f3e3-261b-4320-9e2d-0ed0cfd45906 · outbound

This paper cites Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Theregretful agent: Heuristic-aided navigation through progress estimation, in: CVPR, pp

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:13.181081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.204524Z digest=sha256:6b3a12d6b830e1c030d33c1cc8dd5c84d9cdd8c3afb996d7f17159460a5d12c6

Observation 0187b010-9da3-4203-82d9-cbc16ca6e155 · outbound

This paper cites Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Improving vision-and-language navigation with image-text pairs from the web, in: ECCV, pp

Reference 42

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:12:14.492370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.208375Z digest=sha256:420fcd3edd76aca1154d5fc681e7fc5e2b9013c7bb677def7b452e536f2d5cb6

Observation d50d4463-a269-43cd-9579-b13975f35f18 · outbound

This paper cites SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations SOAT: A scene- and object-aware transformer for vision-and-language navigation, in: NeurIPS, pp

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.478509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.212340Z digest=sha256:354bf0379d267e4b26caa4f43569f2f6f1fba1e50b9a7eac7a4649f12b265695

Observation 9396e758-f39a-412f-bd9a-ccae3a5eaafa · outbound

This paper cites Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bridging the visual semantic gapinVLNviasemanticallyricherinstructions,in:ECCV,pp.54–69

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.220240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.220240Z digest=sha256:cbfc0640bf7c4d1c790e69e724525dbe47a4e9c03ed9510bc12140efa47d4c23

Observation b2052fa8-2f1e-4b64-8f95-d40c067ac602 · outbound

This paper cites Teach: Task-driven embodied agents that chat, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Teach: Task-driven embodied agents that chat, in: AAAI, pp

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.450898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.224781Z digest=sha256:b4dded797dce7df97727e58981ef16fc710adb1d271baa92fdb530849a7c87e2

Observation 88509fdc-81c0-4ebc-bac5-551fa243ac51 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Bleu: a method for automatic evaluation of machine translation, in: Proceedings of the 40th Annual Meeting of the Association for Computational Lin- guistics, pp

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.235094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.235094Z digest=sha256:217c615485760f2d0224391b27e99f35627a7aaaa8fa96151ef9ec5316c47324

Observation 661964f5-77b7-4b2e-9318-550233c04d4a · outbound

This paper cites The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations The road to know-where: An object-and-room informed sequential BERT for indoor vision-language navigation, in: ICCV, pp

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.436917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.239091Z digest=sha256:249e0357b2be040a0fa2c024fad21f50063a4c2f1fd1f3ecb4532302259441f4

Observation 23b9ec66-062c-4738-82ce-742964fb1c6c · outbound

This paper cites Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Object- and-actionawaremodelforvisuallanguagenavigation,in:ECCV,pp

Reference 48

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.419946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.247499Z digest=sha256:7444fc9bc4cb49d5dcd89c8bc785afa32583d2b2d333f53ab0528704e879789b

Observation 2c28b650-edb2-45fc-bfd5-f82a69f624fe · outbound

This paper cites HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations HOP+: history-enhanced and order-aware pre-training for vision- and-language navigation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.256573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.256573Z digest=sha256:9faf890826362e1e25feaf3f43e6532d11d5ea1eae582f64f085a5a79756cccf

Observation 615a9ec6-fb64-4bf0-b055-d42276909213 · outbound

This paper cites Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learningtransferablevisualmodelsfromnatural language supervision, in: ICML, pp

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.422608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.260608Z digest=sha256:5d2fd8ae93f4f262e6d38f6a5a03fcce8492db0ab822672731b99e91debaa7aa

Observation 0fea42e3-ee92-4f61-9f2a-ff3cf639a98c · outbound

This paper cites Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Towards long-horizon vision-language navigation: Platform, benchmark and method, in: CVPR

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.407526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.268752Z digest=sha256:1f52ba4ce1371d954b46febf562e4bcd89e73b471ce1d7168d1c670561082c79

Observation bb77a352-9934-4565-ae53-20ebda4ace23 · outbound

This paper cites Contrastive search is what you need for neuraltextgeneration.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Contrastive search is what you need for neuraltextgeneration

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.393190Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.272856Z digest=sha256:c0e419a813a63f47d575fc4f1838c77f77f236a2762d7b465176b73d1c64bac0

Observation aecb1bbf-1e92-4d93-933c-7851a130c470 · outbound

This paper cites Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Learning to navigate unseen envi- ronments:Backtranslationwithenvironmentaldropout,in:NAACL- HLT, pp

Reference 55

Resolution
verified exact
doi, observed 2026-08-07T05:12:12.405485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.277578Z digest=sha256:9a7d9ce4052949e99b4eff4363495d1a945fe2753f269c48ae4a9aab145b424a

Observation 813c2439-0122-4ab6-935c-edc2581e9a80 · outbound

This paper cites Vision-and-dialog navigation, in: CoRL, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-and-dialog navigation, in: CoRL, pp

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.378885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.281829Z digest=sha256:e253cb9c238f809e9ae7c66e0d9c3cbb0a6a666262616a53b8c382f1c4945961

Observation fa53d099-7623-4a10-9e7a-de62d4e3d867 · outbound

This paper cites Cider: Consensus- based image description evaluation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Cider: Consensus- based image description evaluation, in: CVPR, pp

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.285974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.285974Z digest=sha256:5fed213e40e83137d0d3895c7261ab44d442a692088cd4b74d1188fc17f6b50c

Observation 04a3101a-bf63-467a-b49b-d5e2dec4bf35 · outbound

This paper cites Show and tell: A neural image caption generator, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show and tell: A neural image caption generator, in: CVPR, pp

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.290006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.290006Z digest=sha256:6fe6c9e1ff91c1026f12df57bf985566d09d16657b29f4a90e5cba887a6ffa9b

Observation 3ab685b3-0fb2-443e-b9da-57a4ddd0134d · outbound

This paper cites Soft expert reward learning for vision-and-language navigation, in: ECCV, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Soft expert reward learning for vision-and-language navigation, in: ECCV, pp

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.363551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.298071Z digest=sha256:9ec0556cde73fa168a7da832aad2ee72dd2a84e6f8c081a56a414f0702d8f2e0

Observation 225f25e8-056b-4056-9d25-f6a7a05c2bf7 · outbound

This paper cites GIT: A generative image-to-text transformer for vision and language.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations GIT: A generative image-to-text transformer for vision and language

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.347964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.303260Z digest=sha256:3c3d021a3b6fd317a6bc02d30f07fe4d7d48dc88988ba7df68ae0990d27e6ff5

Observation 341561be-031f-4fea-af60-01fe927b83a4 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.320063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.311490Z digest=sha256:5db66ef77f5145e1ebfe66861ba2846a27f67b71fa869a21b11790b10c53fcfb

Observation 63c57a3a-3ba8-4d14-8b8f-16d07ccc25f9 · outbound

This paper cites OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations OFA: unifying architectures, tasks, and modalities through a simple sequence-to-sequence learning frame- work, in: ICML, pp

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.306372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.319606Z digest=sha256:9a872ee265db9bb667fbd22bef13e6833711272a76fa8e575dd4aaf25eeb4930

Observation 0609499b-c5eb-4540-b01f-9e7fb8b15770 · outbound

This paper cites Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Less is more: Generating grounded navigation instructions from landmarks, in: CVPR, pp

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.323819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.323819Z digest=sha256:38ac6324dbca09781c2398cc524dfb7e5434d3aa413d834fe935ce5064e4e25f

Observation 9edd9f04-e220-4576-822d-f24e52c0f593 · outbound

This paper cites Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Reinforced cross-modal matching and self-supervised imitation learning for vision-language navigation, in: CVPR, pp

Reference 65

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.756522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.327661Z digest=sha256:9e45fea13075ce6ab54e241c9e2bc02765e63138570bcb418dd942696f390e82

Observation 34bea04b-a9f1-4154-b860-46b8db0f2882 · outbound

This paper cites LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations LANA: A language- capablenavigatorforinstructionfollowingandgeneration,in:CVPR, IEEE

Reference 66

Resolution
verified exact
raw_fallback, observed 2026-08-07T05:12:12.677546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.331590Z digest=sha256:99ff35be27d8384440244a7cb71e30709259999546dcc8fc14f85a781c0be160

Observation 3ac4d94b-438e-41ff-acfd-bbfae8db2b06 · outbound

This paper cites Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Show, attend and tell: Neural image :Preprint submitted to Elsevier Page 14 of 15 caption generation with visual attention, in: ICML, pp

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:12:14.292112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.336239Z digest=sha256:866921a6305b2eb38ebfc2a4aed0b88270c1840ecb8231abe525e08e0cb3c58b

Observation d4d4cf0d-a8ac-4376-9cf1-1a02599e6059 · outbound

This paper cites Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unified vision-language pre-training for image captioning and VQA, in: AAAI, pp

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.340247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.340247Z digest=sha256:a1b90353d108a957d1b06c0319fe77597906385835a7d00d0ebc8836b61f6df8

Observation 6babf70e-4066-4e21-8f4c-c977d6cd47cd · outbound

This paper cites Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Vision-language navigation with self-supervised auxiliary reasoning tasks, in: CVPR, pp

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.345285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.345285Z digest=sha256:47a87ddfb85c4f767590d2fa89e7a68661325464ff6e6b65c5cd1f29ca2ece58

Observation ca6b9874-9a29-4ed5-86d9-409884f01816 · outbound

This paper cites Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Babywalk:Goingfartherinvision-and-languagenavigationbytaking babysteps,in:ACL,pp.2539–2556

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.349149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.349149Z digest=sha256:ba59002977c4ab294467e3233bc3b19180ee6cea8f9d2e43a158b737081a4a1e

Observation 026645a0-97a9-4a3b-9725-f08d96dcf612 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 380

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.085959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.085959Z digest=sha256:8a203003e8246bb18328da64439e73b496cdf2e7cc19d079801e0a7e8eace595

Observation b910639c-c43c-4e4b-b389-5cbba8e5a53d · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 676

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.045804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.045804Z digest=sha256:7c066d6d8916f69b18da4461972e7629aeeda0f28c804baeeabfcf78144d1377

Observation cf621fab-ba65-4bc6-a139-d8eb16438d01 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 1644

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.243135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.243135Z digest=sha256:ed1baf2f48a80beddab465e73cd42fe9a7495b1db67147cd2f0dd6e3b67d21d7

Observation 538109e6-4675-4359-a24c-0044a6b7a330 · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.333845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.307332Z digest=sha256:73cf6f56095e10459b9e2ac31e65ed63c9c9103ee3e90640ce20c43c8a125791

Observation 160a8a3a-27e1-40c4-bc54-ab34e3c7420b · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2024

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:12:12.830343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.315636Z digest=sha256:324afb9e1194750453e15960efc355f8e694f250761d3a1ff688db7bd9766dc3

Observation 56374228-da84-4be0-885b-62f890fe8cef · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-07T05:12:12.229720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:12:12.229720Z digest=sha256:e69bfcdba50cdf7d0d2886968dd3c36a0826998e67893c2dbbd3c245926af53f

Observation 8fd95b62-9716-4c42-8769-55257bce25fb · outbound

This paper cites an unresolved cited work.

Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations Unresolved cited work

Reference 7367

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:12:14.465023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:12:12.216264Z digest=sha256:697e7d2d6ec763e039a9696f3977447048a834438256e63b43c38146ab36e693

Pith citing papers

Observation 134c4ef6-5447-4a21-bc5c-649dbd55a063 · inbound

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning cites this paper.

GoViG: Goal-Conditioned Visual Navigation Instruction Generation via Multimodal Reasoning Generating Vision-Language Navigation Instructions Incorporated Fine-Grained Alignment Annotations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-18T23:02:52.670999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T23:02:14.913189Z digest=sha256:5f4166e84a444076891071b1eaa93db34166fec7ebcca087461836fb3f4390ee