Pith. sign in

Paper Citation Record · LEDGER

What is needed for simple spatial language capabilities in VQA?

As of 16 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:1908.06336.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.06336 v2

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:51:50.130298Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1270a9a5-ce52-44d0-8fd8-c72545c7c2b5 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

What is needed for simple spatial language capabilities in VQA? Lawrence Zitnick, and Devi Parikh

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.272851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.088573Z digest=sha256:b7c198a3e267ecb64f062e0250be2ca2a3b4edd84ad2bb58adffd4b97ee58a48

Observation a3c5d07f-69ba-449c-9cd3-81d42d2e5db2 · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering.

What is needed for simple spatial language capabilities in VQA? Learning to reason: End-to-end module networks for visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.263901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.092487Z digest=sha256:68e1355ae579b11f63623478538ecc5f67719eae2edff445b0d752d208b3f26a

Observation debc4da1-3e7d-4eea-b2d2-a60942b68be2 · outbound

This paper cites Hudson and Christopher D.

What is needed for simple spatial language capabilities in VQA? Hudson and Christopher D

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.254433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.096183Z digest=sha256:d7a78279a0678fa60c22a0e35ce27140dc2df80de20f07d5def66a25e8dbc5d5

Observation 034be91e-4d21-4ee4-8866-68b66cf0a199 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

What is needed for simple spatial language capabilities in VQA? Lawrence Zitnick, and Ross Girshick

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.246041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.099683Z digest=sha256:71a8e93ce39125e831b27263ee1382b80020148820a73d10deb19026840e25d3

Observation bdb221b6-0c17-43c2-ac4e-7ad356f183be · outbound

This paper cites Inferring and executing programs for visual rea- soning.

What is needed for simple spatial language capabilities in VQA? Inferring and executing programs for visual rea- soning

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.236688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.103235Z digest=sha256:41711d3eac9ada8f9928dfe88b132770ebee8931bf7044d9e1ec335b55689cf5

Observation b5f8e9fd-5729-4509-9b38-a0b31798e70a · outbound

This paper cites ShapeWorld - A new test methodology for multimodal language understanding.

What is needed for simple spatial language capabilities in VQA? ShapeWorld - A new test methodology for multimodal language understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:51:50.106637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:51:50.106637Z digest=sha256:168c820dbb9b30dc81842f34bc6327bda8e5adde41f5e4be8e6869e34a1e6eb9

Observation 34d41f2c-1ef3-4c03-965a-9f685886a2c1 · outbound

This paper cites An intriguing failing of convolutional neural networks and the CoordConv solution.

What is needed for simple spatial language capabilities in VQA? An intriguing failing of convolutional neural networks and the CoordConv solution

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.227187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.110709Z digest=sha256:5dd0e1d2a9abf44eef6c0e3b5358823c0555ee337bc9a0020aa4f909962214a2

Observation 2ff7f002-23bf-495c-8b9f-7ca4293b3dad · outbound

This paper cites The visual QA devil in the details: The impact of early fusion and batch norm on CLEVR.

What is needed for simple spatial language capabilities in VQA? The visual QA devil in the details: The impact of early fusion and batch norm on CLEVR

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.216514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.113934Z digest=sha256:2ab96b723d8c6d900f3317cef7797a167941dccc23fa8857d0ddb56c7d36eaf8

Observation cced3cac-9939-4f2a-bb84-29d3de31aa0a · outbound

This paper cites Transparency by design: Closing the gap between performance and interpretability in visual reasoning.

What is needed for simple spatial language capabilities in VQA? Transparency by design: Closing the gap between performance and interpretability in visual reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.206412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.117280Z digest=sha256:bc0911ce9cab9f6ba4e67aac21a4eb99b4387abeed93b614bcd15e232253f2df

Observation 537d91c6-afb8-477c-b3df-dadc5de4c03c · outbound

This paper cites Courville.

What is needed for simple spatial language capabilities in VQA? Courville

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.196498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.120707Z digest=sha256:941b77ae2fc5d19b8ce9982179c1bb4221e5a842d36364588b1e1e295e40908c

Observation 9f6edb19-3c55-42d4-ba50-9fb83c13623e · outbound

This paper cites an unresolved cited work.

What is needed for simple spatial language capabilities in VQA? Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:51:50.186887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.123903Z digest=sha256:1ff6527fc600d11d624c4c2795a79d7effa15136263c835ff19143772863d4e5

Observation c549683c-4f8f-4904-8b57-9ac304deffd7 · outbound

This paper cites A dataset and architecture for visual reasoning with a working memory.

What is needed for simple spatial language capabilities in VQA? A dataset and architecture for visual reasoning with a working memory

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T12:51:50.177406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.126987Z digest=sha256:883e8f191384b8111a92aecb90aa0f47d6ee81348d53f32c428f5d2765d9c6ed

Observation c2cc1188-cf79-4913-8f0a-02345e2d1038 · outbound

This paper cites an unresolved cited work.

What is needed for simple spatial language capabilities in VQA? Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:51:50.166533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-14T12:51:50.130298Z digest=sha256:24c6dbf86804f32a24b7ce875f9eebc516bc402e5d407306552dddc7a7ef5e7b

Pith citing papers

No inbound Pith citation observations are available.