Pith. sign in

Paper Citation Record · LEDGER

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering

As of 15 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:1908.04950.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04950 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:32:25.742010Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact1
  • verified fuzzy23
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb66dcb0-3718-4272-97f1-4862dde0b7a5 · outbound

This paper cites Blindfold Baselines for Embodied QA.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Blindfold Baselines for Embodied QA

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.392507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.392507Z digest=sha256:97517f8ea5fef97602c49dcdf7f2da16a31e1b99e916cecf45e81980085d20ed

Observation 3ceb717c-cb95-4285-9818-7c304bbd45d0 · outbound

This paper cites VQA: Visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering VQA: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.161905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.403046Z digest=sha256:75391339a11caf20ceef96d54cff1e6394de760d4b9ad1aa2a6c9b936ffabb91

Observation bb7cb70d-7249-41d5-b63f-432c24679e83 · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Machine Translation by Jointly Learning to Align and Translate

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.416278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.416278Z digest=sha256:57f969371e86269381032abefc90a7a11b6dd8e73d8989602231c545f8214a68

Observation f698ce84-25f4-4e73-a082-099eaa098921 · outbound

This paper cites Systematic Generalization: What Is Required and Can It Be Learned?.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Systematic Generalization: What Is Required and Can It Be Learned?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.423479Z digest=sha256:24bbeab0bb9ee63e267575ba15a756f55028286c0ec5e6b587e23f67b457b3f8

Observation d9435600-09ec-4fb2-b0af-0f1bf2f7a7e5 · outbound

This paper cites MUREL: Multimodal Relational Reasoning for Visual Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MUREL: Multimodal Relational Reasoning for Visual Question Answering

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.133297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.432553Z digest=sha256:156f522d533dfdd25bd5f0def52c52deba3af965dfcf1bf47e6de498e07cf968

Observation 21037c04-3690-42e4-9a09-ab511b70b655 · outbound

This paper cites Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning Phrase Representations using RNN Encoder-Decoder for Statistical Machine Translation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.441540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.441540Z digest=sha256:d56ff90105f761bd40c6aa81f533111ccd80d7ab7ee838ccba53fcf9fc2f089a

Observation d902790f-c848-4070-b491-26252b6a3801 · outbound

This paper cites Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.106821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.449561Z digest=sha256:2dbaaff0d0b1bfaae7f8f7eae0c50cc61ffaa8b94fc9febb987099ad1ab822e8

Observation 0611e946-1323-42c6-b858-d789e6676321 · outbound

This paper cites Neural Modular Control for Embodied Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Neural Modular Control for Embodied Question Answering

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.457445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.457445Z digest=sha256:9d45b80959842812c72feb963b2c27d4aeb39363be04536b101a63cb13da7d3f

Observation e7281b40-acab-49ad-ad69-0787fb0744d3 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.464101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.464101Z digest=sha256:73e1ca996697c489051a1a79aeb6220a3a93be1604262d0113914c684516b8a8

Observation c4b69e52-cbce-46cf-99b2-be8ad2726443 · outbound

This paper cites Deep sparse rectifier neural net- works.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Deep sparse rectifier neural net- works

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.084683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.469712Z digest=sha256:9e42c82c3feab60faebfcd7c2fb22d94977c8e06be7e96cb659b0c1194967244

Observation 25ae027a-3402-45cb-bb46-353cc3a5dae2 · outbound

This paper cites IQA: Visual question answering in interactive environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering IQA: Visual question answering in interactive environments

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.054516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.475855Z digest=sha256:63c419680a7d31b339ce2b8eb4ba2bcf2122e68507dfb6e848841eebe28ac8af

Observation 580fc635-eca8-4468-a0ab-dde7ab6ca1f2 · outbound

This paper cites Long short-term memory.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Long short-term memory

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.481520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.481520Z digest=sha256:24971fc286be9b10f1efcc2cb103714dba002ef8f7fad30ce1a41d1304837984

Observation 4c4f30e5-db58-4caf-9271-0de752a17aba · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning to reason: End-to-end module networks for visual question answering

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:27.007765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.489048Z digest=sha256:f45404a339f270b84c7393784868fb70ba4824d8c60ac1c205c26b91c62da6d3

Observation 8668e006-e0f2-42ff-a27d-75aa378c41fd · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Compositional Attention Networks for Machine Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.498398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.498398Z digest=sha256:249ca3a9670ce8e85c856086db72524419cfdba2a588919b238d3942b16cf37c

Observation 5d8ca0f3-662c-475c-9429-cf4b9f0a345d · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.507069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.507069Z digest=sha256:2ec780ee031bb90db8cdd52451f806fdb4940d7b556a7e22cdae066845c8e64c

Observation 2f4ecd35-991c-4f79-838f-d2be5c8057e6 · outbound

This paper cites Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.514240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.514240Z digest=sha256:0392bb92719965d02928d43b863ac9f57250bfe46abb46f8c0ca11baec1c01c4

Observation aac89881-1bbe-4c80-97d8-1b4d1b361ced · outbound

This paper cites CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering CLEVR: A diagnostic dataset for compo- sitional language and elementary visual reasoning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.983480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.524485Z digest=sha256:49f89ebf7f4985bb6f37dad8cd3705f10f530bf3f41e2eae37f9bdb49e78fd50

Observation 5667ffb2-fe37-4b0f-9b30-5dedac342b26 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Adam: A Method for Stochastic Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.531651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.531651Z digest=sha256:94bc0cbb14af29b630aba3da8c5c7ca1a407b1269e881b1704193e72a94dfd80

Observation 8c93cb0d-202b-46c5-a396-3ae3e9b56257 · outbound

This paper cites AI2-THOR: An Interactive 3D Environment for Visual AI.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering AI2-THOR: An Interactive 3D Environment for Visual AI

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.957511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.544151Z digest=sha256:a65c3647ba932c6dbe530507eb47ee41e0be02b722c6dd90e968b2882c9cf962

Observation a5d61e02-1473-4008-9229-22968a93c996 · outbound

This paper cites TVQA: Localized, Composi- tional Video Question Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering TVQA: Localized, Composi- tional Video Question Answering

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.924883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.552668Z digest=sha256:0f25a648f1abfea4101e39ebbb742c22ddc79aacf787da6b46184c5a5fbcdd35

Observation 5923f293-cfe4-4b89-a1ea-0378f6ec0dc3 · outbound

This paper cites Learning visual question answering by bootstrapping hard attention.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning visual question answering by bootstrapping hard attention

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.890519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.560966Z digest=sha256:add62f9ee36f27cb8bd4b0ee5e0f94653ae19262a8ef8c45d42f0d6329907ba7

Observation 4662b04d-38d0-4ce2-a46a-e98bae3fb0ee · outbound

This paper cites Benchmarking Classic and Learned Navigation in Complex 3D Environments.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Benchmarking Classic and Learned Navigation in Complex 3D Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.568564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.568564Z digest=sha256:0af0ace451e13e8adad444556fdab73a00babb04be5b7d3c3896b9a5f84a516d

Observation dc25097c-b4e1-4eab-85ef-b28af294b1f3 · outbound

This paper cites MarioQA: An- swering Questions by Watching Gameplay Videos.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MarioQA: An- swering Questions by Watching Gameplay Videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.867007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.583764Z digest=sha256:46521f4d1bfb59d5d22dd75444d8957e57b0f6f50e75d08ed3b8bb877ad551f1

Observation 07c20715-2e7c-4fac-870e-9046ed3ed624 · outbound

This paper cites Out of the box: Reasoning with graph convolution nets for factual visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Out of the box: Reasoning with graph convolution nets for factual visual question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.825665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.596086Z digest=sha256:bf1df71f62d3db40da44febfecf3bb6377a1f564ef447f6463022961c068ac81

Observation 96e187f4-df2d-44e5-a2a6-c47974ab5ccb · outbound

This paper cites From FiLM to Video: Multi-turn Question Answering with Multi-modal Context.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering From FiLM to Video: Multi-turn Question Answering with Multi-modal Context

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:32:25.943888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.601770Z digest=sha256:29431fc9b5cbe30682ac28fc21a18d01c337d9ccf7f3457fab002f37d7b8ce22

Observation 7774f0a4-0872-4422-a644-259a5252035c · outbound

This paper cites Learning conditioned graph structures for interpretable visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning conditioned graph structures for interpretable visual question answering

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.789387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.609549Z digest=sha256:a06adb731baf378f03393347fb1d94a529e054f13998a7c91306e91c509e687b

Observation fd108471-39a0-4854-aa2d-fd978956a179 · outbound

This paper cites FiLM: Visual reasoning with a general conditioning layer.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering FiLM: Visual reasoning with a general conditioning layer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.759038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.616696Z digest=sha256:5a3a28027ec3e65a6ad0c9443bae7aa013e44071ce24233c592c6fbd7983d2d5

Observation 9b2716db-af14-4eb8-b7d3-59523db26a58 · outbound

This paper cites Exploring models and data for im- age question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Exploring models and data for im- age question answering

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.730829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.625152Z digest=sha256:b24803e8ab433537e5d755e0891552c094e7298391db38c3f006336827cf08f1

Observation 4034e400-0c87-4e4e-ab0b-4fb1589ae54e · outbound

This paper cites Faster R-CNN: Towards real- time object detection with region proposal networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Faster R-CNN: Towards real- time object detection with region proposal networks

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.698989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.632208Z digest=sha256:9235dcbaddc2d83fa91e3daf69d40721f552cef210a6233851e7d40b44ad3499

Observation d79b820c-bf94-4b62-b008-a4ae3d148811 · outbound

This paper cites Habitat: A Platform for Embodied AI Research.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Habitat: A Platform for Embodied AI Research

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.640466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.640466Z digest=sha256:3c4fd271f8c73e23ce29171944b9a5c6842fdfed0151132c8be6fc581569f502

Observation 609d4e1c-5610-49f8-b2b0-dd44ac332433 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.650423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.650423Z digest=sha256:e96d49ff2551451aba1bc575750da0bb00c8d331fc2169349303405611eb4ec3

Observation 45dde7b6-6029-4324-8100-1548eac231b3 · outbound

This paper cites Semantic Scene Completion from a Single Depth Image.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Semantic Scene Completion from a Single Depth Image

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.651023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.658370Z digest=sha256:b79e8a0b38be41b2155d87eff882ee5c7170ee858588c5e07c5efd3b485dc15f

Observation 9c4266da-d693-4479-aa03-dfee70061caf · outbound

This paper cites Visual Reasoning with Multi-hop Feature Modulation.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Visual Reasoning with Multi-hop Feature Modulation

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.607478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.672201Z digest=sha256:c49b40baa6597db0121dad22324ae0e8a83f8dee683dea135c5ccdd721f86af9

Observation 968ea279-622b-4d55-8611-1bfe24ce2d6c · outbound

This paper cites MovieQA: Understanding Stories in Movies through Question- Answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering MovieQA: Understanding Stories in Movies through Question- Answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.579054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.681109Z digest=sha256:bf0a18dca0257c8a51a05cd8b6f4616f450b74e1912f1f57d9a38680763fe31a

Observation 77497a4e-c0b6-4329-af56-0812a0ce3c07 · outbound

This paper cites Graph-structured repre- sentations for visual question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Graph-structured repre- sentations for visual question answering

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.546012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.691890Z digest=sha256:3c04cb40ad0a00ea4e92172c46c679dfda0f16a6e98f96a648e3d886642b81ac

Observation c9a2c48c-af4e-4097-ba01-88cb44e7b29d · outbound

This paper cites Learning spatiotemporal features with 3D convolutional networks.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Learning spatiotemporal features with 3D convolutional networks

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.515729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.709462Z digest=sha256:a64b20c008ae6018f58317c68ef880751c5e27dbe64a114752794f3b46b9d441

Observation e0f33697-411a-4f17-815f-210133603017 · outbound

This paper cites Embodied Question Answering in Photorealistic Environments with Point Cloud Perception.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Embodied Question Answering in Photorealistic Environments with Point Cloud Perception

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.483075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.716852Z digest=sha256:4022950f25aa0f7ce343b9f8ce5af73a4f8f155f295d6b1da5b296a19cbfad0b

Observation ee8e3cb8-e2a5-4a54-9dec-db60094833bc · outbound

This paper cites Building Generalizable Agents with a Realistic and Rich 3D Environment.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Building Generalizable Agents with a Realistic and Rich 3D Environment

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.726025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.726025Z digest=sha256:7868fb6298d3ba0c9592e94cae74afdb16bdef5de2598acf51f94ec0f2c76919

Observation b8195918-9c8b-482b-a52c-36c9ccd4362f · outbound

This paper cites Stacked attention networks for image question answering.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Stacked attention networks for image question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-14T13:32:25.733408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:32:25.733408Z digest=sha256:c38e520c175ed3efb582fd6356fca412882c58e11d82425ef476649b6164b6d0

Observation 416e6c67-7bb7-461b-8dc7-a15cb969a60d · outbound

This paper cites Berg, and Dhruv Batra.

VideoNavQA: Bridging the Gap between Visual and Embodied Question Answering Berg, and Dhruv Batra

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:32:26.425804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T13:32:25.742010Z digest=sha256:ced47ba132c95d83ea4823882fc65d1a7275ab2e0a6c233273ff2012c75958ac

Pith citing papers

No inbound Pith citation observations are available.