Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:32:25.742010Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:1908.04950.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:32:25.742010Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
40 of 40 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb66dcb0-3718-4272-97f1-4862dde0b7a5 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Blindfold Baselines for Embodied QA
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ceb717c-cb95-4285-9818-7c304bbd45d0 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering VQA: Visual question answering
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation bb7cb70d-7249-41d5-b63f-432c24679e83 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Machine Translation by Jointly Learning to Align and Translate
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f698ce84-25f4-4e73-a082-099eaa098921 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Systematic Generalization: What Is Required and Can It Be Learned?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9435600-09ec-4fb2-b0af-0f1bf2f7a7e5 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MUREL: Multimodal Relational Reasoning for Visual Question Answering
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 21037c04-3690-42e4-9a09-ab511b70b655 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d902790f-c848-4070-b491-26252b6a3801 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0611e946-1323-42c6-b858-d789e6676321 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Modular Control for Embodied Question Answering
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7281b40-acab-49ad-ad69-0787fb0744d3 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b69e52-cbce-46cf-99b2-be8ad2726443 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Deep sparse rectifier neural net- works
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 25ae027a-3402-45cb-bb46-353cc3a5dae2 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering IQA: Visual question answering in interactive environments
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 580fc635-eca8-4468-a0ab-dde7ab6ca1f2 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Long short-term memory
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c4f30e5-db58-4caf-9271-0de752a17aba · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning to reason: End-to-end module networks for visual question answering
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8668e006-e0f2-42ff-a27d-75aa378c41fd · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Compositional Attention Networks for Machine Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d8ca0f3-662c-475c-9429-cf4b9f0a345d · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f4ecd35-991c-4f79-838f-d2be5c8057e6 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aac89881-1bbe-4c80-97d8-1b4d1b361ced · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5667ffb2-fe37-4b0f-9b30-5dedac342b26 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Adam: A Method for Stochastic Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c93cb0d-202b-46c5-a396-3ae3e9b56257 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering AI2-THOR: An Interactive 3D Environment for Visual AI
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation a5d61e02-1473-4008-9229-22968a93c996 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering TVQA: Localized, Composi- tional Video Question Answering
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5923f293-cfe4-4b89-a1ea-0378f6ec0dc3 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning visual question answering by bootstrapping hard attention
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4662b04d-38d0-4ce2-a46a-e98bae3fb0ee · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Benchmarking Classic and Learned Navigation in Complex 3D Environments
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc25097c-b4e1-4eab-85ef-b28af294b1f3 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MarioQA: An- swering Questions by Watching Gameplay Videos
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 07c20715-2e7c-4fac-870e-9046ed3ed624 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Out of the box: Reasoning with graph convolution nets for factual visual question answering
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 96e187f4-df2d-44e5-a2a6-c47974ab5ccb · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering From FiLM to Video: Multi-turn Question Answering with Multi-modal Context
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 7774f0a4-0872-4422-a644-259a5252035c · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning conditioned graph structures for interpretable visual question answering
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd108471-39a0-4854-aa2d-fd978956a179 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering FiLM: Visual reasoning with a general conditioning layer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9b2716db-af14-4eb8-b7d3-59523db26a58 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Exploring models and data for im- age question answering
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4034e400-0c87-4e4e-ab0b-4fb1589ae54e · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Faster R-CNN: Towards real- time object detection with region proposal networks
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d79b820c-bf94-4b62-b008-a4ae3d148811 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Habitat: A Platform for Embodied AI Research
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 609d4e1c-5610-49f8-b2b0-dd44ac332433 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45dde7b6-6029-4324-8100-1548eac231b3 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Semantic Scene Completion from a Single Depth Image
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 9c4266da-d693-4479-aa03-dfee70061caf · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Visual Reasoning with Multi-hop Feature Modulation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 968ea279-622b-4d55-8611-1bfe24ce2d6c · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MovieQA: Understanding Stories in Movies through Question- Answering
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 77497a4e-c0b6-4329-af56-0812a0ce3c07 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Graph-structured repre- sentations for visual question answering
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c9a2c48c-af4e-4097-ba01-88cb44e7b29d · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning spatiotemporal features with 3D convolutional networks
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation e0f33697-411a-4f17-815f-210133603017 · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering in Photorealistic Environments with Point Cloud Perception
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ee8e3cb8-e2a5-4a54-9dec-db60094833bc · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Building Generalizable Agents with a Realistic and Rich 3D Environment
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8195918-9c8b-482b-a52c-36c9ccd4362f · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Stacked attention networks for image question answering
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 416e6c67-7bb7-461b-8dc7-a15cb969a60d · outbound
VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Berg, and Dhruv Batra
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
No inbound Pith citation observations are available.