Pith. sign in

Paper Citation Record · LEDGER

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering

As of 14 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:1908.04950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04950 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:32:25.742010Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb66dcb0-3718-4272-97f1-4862dde0b7a5 · outbound

This paper cites Blindfold Baselines for Embodied QA.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Blindfold Baselines for Embodied QA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.392507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.392507Z digest=sha256:f47063ab441f345d8b296ccbc830c6447952f16e648c885e61ac29b6b14a382e

Observation 3ceb717c-cb95-4285-9818-7c304bbd45d0 · outbound

This paper cites VQA: Visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering VQA: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.161905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.403046Z digest=sha256:404ab7a5de21b4bde90694ce51d7606e688920df70025ff2f42997e8251f741b

Observation bb7cb70d-7249-41d5-b63f-432c24679e83 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Machine Translation by Jointly Learning to Align and Translate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.416278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.416278Z digest=sha256:8095a59471c36a6cf3b476de722b8263813c4a7473ca05d43202af69543b0013

Observation f698ce84-25f4-4e73-a082-099eaa098921 · outbound

This paper cites Systematic Generalization: What Is Required and Can It Be Learned?.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Systematic Generalization: What Is Required and Can It Be Learned?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.423479Z digest=sha256:76724639567597968150777c6917da60f9a94ffa7727936e1c9c88c0bf123512

Observation d9435600-09ec-4fb2-b0af-0f1bf2f7a7e5 · outbound

This paper cites MUREL: Multimodal Relational Reasoning for Visual Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MUREL: Multimodal Relational Reasoning for Visual Question Answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.133297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.432553Z digest=sha256:88ce860b70b4349b8440e4d472d45724d161e594a237b02210da3047ce245f81

Observation 21037c04-3690-42e4-9a09-ab511b70b655 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.441540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.441540Z digest=sha256:b1cd536291454430093bafd104da516f33c4327ab08c5911403c90c64c19479f

Observation d902790f-c848-4070-b491-26252b6a3801 · outbound

This paper cites Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.106821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.449561Z digest=sha256:b86779fd27de1220d3adf1b4ce8b89d23143d7381d97ecfe417a4826df2838ee

Observation 0611e946-1323-42c6-b858-d789e6676321 · outbound

This paper cites Neural Modular Control for Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Modular Control for Embodied Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.457445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.457445Z digest=sha256:250b4c364487b8e8521051a067b9719ecbaf1ed87b1f2b078c302ca3b5354f20

Observation e7281b40-acab-49ad-ad69-0787fb0744d3 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.464101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.464101Z digest=sha256:31a2edc9dbf7292fd1c868ec129bbcd56fb362007f5ff6b79d3816734810ced4

Observation c4b69e52-cbce-46cf-99b2-be8ad2726443 · outbound

This paper cites Deep sparse rectifier neural net- works.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Deep sparse rectifier neural net- works

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.084683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.469712Z digest=sha256:c768524fb4e13a72212dbf4a8dcf7a64d8e742892ef0dac7d10470fabb1dfe6a

Observation 25ae027a-3402-45cb-bb46-353cc3a5dae2 · outbound

This paper cites IQA: Visual question answering in interactive environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering IQA: Visual question answering in interactive environments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.054516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.475855Z digest=sha256:431c8736706826d5340676f6060ca4e8dd9fd3f05493e85fdff27afe6de8169e

Observation 580fc635-eca8-4468-a0ab-dde7ab6ca1f2 · outbound

This paper cites Long short-term memory.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Long short-term memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.481520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.481520Z digest=sha256:2ac167afe2f81bb017675e702bb3d489c9045917e3c142e79454bcbe1247cfc9

Observation 4c4f30e5-db58-4caf-9271-0de752a17aba · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning to reason: End-to-end module networks for visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.007765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.489048Z digest=sha256:e702dccbe89f802cb3a580f222efb569222044324889983d24e1edd830f76ab2

Observation 8668e006-e0f2-42ff-a27d-75aa378c41fd · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Compositional Attention Networks for Machine Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.498398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.498398Z digest=sha256:07e1125566e1ef23bf3ade15ede91cc61bb76ef69fde14b7df7b4cf534a70c99

Observation 5d8ca0f3-662c-475c-9429-cf4b9f0a345d · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.507069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.507069Z digest=sha256:6a8860a92df5c6d29a5b450b2f35694e8545fc5d065965b570302a9717737945

Observation 2f4ecd35-991c-4f79-838f-d2be5c8057e6 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.514240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.514240Z digest=sha256:94b0b2c4db1a116f20277b8b529973738818405a241c8298c39083cc5450137a

Observation aac89881-1bbe-4c80-97d8-1b4d1b361ced · outbound

This paper cites CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.983480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.524485Z digest=sha256:5dbf87ee23e266c65909d960dfa5d97978277ed80213a4ebfdbb00e59de87cc1

Observation 5667ffb2-fe37-4b0f-9b30-5dedac342b26 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.531651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.531651Z digest=sha256:b98352b5503ed5c7cd9274c88d85335e9bb636e47b0ce3c6ba75711ce4e565bb

Observation 8c93cb0d-202b-46c5-a396-3ae3e9b56257 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.957511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.544151Z digest=sha256:7e759fb8844f8151b430fe03b15bbed1d0ec66d42d6b9a90c4660a5298b4f1cd

Observation a5d61e02-1473-4008-9229-22968a93c996 · outbound

This paper cites TVQA: Localized, Composi- tional Video Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering TVQA: Localized, Composi- tional Video Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.924883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.552668Z digest=sha256:39197e2125e56c3146704144e6177be3e9ecd6b46fb9f875c293a321127e507e

Observation 5923f293-cfe4-4b89-a1ea-0378f6ec0dc3 · outbound

This paper cites Learning visual question answering by bootstrapping hard attention.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning visual question answering by bootstrapping hard attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.890519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.560966Z digest=sha256:d622b1abaaef7212fa86cea2d847d0ee449de0598d894a3e4f1c1018dc095b27

Observation 4662b04d-38d0-4ce2-a46a-e98bae3fb0ee · outbound

This paper cites Benchmarking Classic and Learned Navigation in Complex 3D Environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Benchmarking Classic and Learned Navigation in Complex 3D Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.568564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.568564Z digest=sha256:5a3a951e5c7f9e950aa6b122bcf6defee21b59150879424765fedf96fad1f1d3

Observation dc25097c-b4e1-4eab-85ef-b28af294b1f3 · outbound

This paper cites MarioQA: An- swering Questions by Watching Gameplay Videos.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MarioQA: An- swering Questions by Watching Gameplay Videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.867007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.583764Z digest=sha256:424e1ee7c4c933752899d38395ba794d613a9dbd4e45ed1e689ae07193062d2f

Observation 07c20715-2e7c-4fac-870e-9046ed3ed624 · outbound

This paper cites Out of the box: Reasoning with graph convolution nets for factual visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Out of the box: Reasoning with graph convolution nets for factual visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.825665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.596086Z digest=sha256:ee8984a001b794d54cdcef392b44e029e254595102947dc96bdaa0d3778e06ec

Observation 96e187f4-df2d-44e5-a2a6-c47974ab5ccb · outbound

This paper cites From FiLM to Video: Multi-turn Question Answering with Multi-modal Context.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering From FiLM to Video: Multi-turn Question Answering with Multi-modal Context

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:32:25.943888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.601770Z digest=sha256:0729c5013bf6cf5e303c5bea2ff6365c88ae248a3f570247b21e38dba23bf1d3

Observation 7774f0a4-0872-4422-a644-259a5252035c · outbound

This paper cites Learning conditioned graph structures for interpretable visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning conditioned graph structures for interpretable visual question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.789387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.609549Z digest=sha256:57d18951560f3c38555c9f108d90d05f858fd9a79eaab722b8a4d814adc78824

Observation fd108471-39a0-4854-aa2d-fd978956a179 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering FiLM: Visual reasoning with a general conditioning layer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.759038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.616696Z digest=sha256:6c42a2148c14550a7a9f90b7e442804922005c2a3ac50abba228257eeafdf910

Observation 9b2716db-af14-4eb8-b7d3-59523db26a58 · outbound

This paper cites Exploring models and data for im- age question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Exploring models and data for im- age question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.730829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.625152Z digest=sha256:ded3b58a63f13fa3a2428c0a9346ae9829affcec184e13f09be8a385635d97ac

Observation 4034e400-0c87-4e4e-ab0b-4fb1589ae54e · outbound

This paper cites Faster R-CNN: Towards real- time object detection with region proposal networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Faster R-CNN: Towards real- time object detection with region proposal networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.698989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.632208Z digest=sha256:243950dff7c0455cb6f6e7248e59524d41f1ea5a4c3babb69427999cc36e7852

Observation d79b820c-bf94-4b62-b008-a4ae3d148811 · outbound

This paper cites Habitat: A Platform for Embodied AI Research.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Habitat: A Platform for Embodied AI Research

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.640466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.640466Z digest=sha256:178286d7b8765cb8fc8fa2aa714679e146012f9ab8823a1e0dc03e86ddca72d1

Observation 609d4e1c-5610-49f8-b2b0-dd44ac332433 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.650423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.650423Z digest=sha256:6b9c3ec13da42e6ec4ff4c1f7b8b396133fbeb8f61123c8ae80fc70b9175b2d1

Observation 45dde7b6-6029-4324-8100-1548eac231b3 · outbound

This paper cites Semantic Scene Completion from a Single Depth Image.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Semantic Scene Completion from a Single Depth Image

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.658370Z digest=sha256:781e80b883b05a5dcc00c1e97740aa2228698814a01fc8c9b353f31053babeee

Observation 9c4266da-d693-4479-aa03-dfee70061caf · outbound

This paper cites Visual Reasoning with Multi-hop Feature Modulation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Visual Reasoning with Multi-hop Feature Modulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.607478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.672201Z digest=sha256:60861838d7b805377b46c9a9830e9268d6dd95934f4365e3804cc7419ba06a80

Observation 968ea279-622b-4d55-8611-1bfe24ce2d6c · outbound

This paper cites MovieQA: Understanding Stories in Movies through Question- Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MovieQA: Understanding Stories in Movies through Question- Answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.579054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.681109Z digest=sha256:94c55ce18dbe7c529cf687b2d4e0f8f4b63c9393c8b424da369db2bc563c9459

Observation 77497a4e-c0b6-4329-af56-0812a0ce3c07 · outbound

This paper cites Graph-structured repre- sentations for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Graph-structured repre- sentations for visual question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.546012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.691890Z digest=sha256:2527158ebc42e203ccfe5270185b4602ce6e6154ae730e6432e9cce0ecb9d56e

Observation c9a2c48c-af4e-4097-ba01-88cb44e7b29d · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning spatiotemporal features with 3D convolutional networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.515729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.709462Z digest=sha256:4f4b8a1d690b0b1145ee9794eee226826de163b5442767e17046910a06470221

Observation e0f33697-411a-4f17-815f-210133603017 · outbound

This paper cites Embodied Question Answering in Photorealistic Environments with Point Cloud Perception.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering in Photorealistic Environments with Point Cloud Perception

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.483075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.716852Z digest=sha256:ff283de504401d2893ca3164fd452ed922fbd2059c11fdd5792bf1d3c9c4e135

Observation ee8e3cb8-e2a5-4a54-9dec-db60094833bc · outbound

This paper cites Building Generalizable Agents with a Realistic and Rich 3D Environment.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Building Generalizable Agents with a Realistic and Rich 3D Environment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.726025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.726025Z digest=sha256:200a3d3a2d519791b5afb362e862937125eca0208b37cdf30c4a90f0b2dc55ba

Observation b8195918-9c8b-482b-a52c-36c9ccd4362f · outbound

This paper cites Stacked attention networks for image question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Stacked attention networks for image question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.733408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.733408Z digest=sha256:132341a1ce505dfdd0cc655dd7ce3565a8dfbefa0aeb6ea7e312250b639cd5b4

Observation 416e6c67-7bb7-461b-8dc7-a15cb969a60d · outbound

This paper cites Berg, and Dhruv Batra.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Berg, and Dhruv Batra

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.425804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.742010Z digest=sha256:2db38a65516939140d154262440f0c937e5f16f4c58ab2a4e555ddafde9a7968

Pith citing papers

No inbound Pith citation observations are available.