Pith. sign in

Paper Citation Record · LEDGER

Fusion of Detected Objects in Text for Visual Question Answering

As of 22 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:1908.05054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.05054 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.432399Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:01:13.232217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:03:58.089358Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ccb0bb9-47f5-490b-9b41-9fbfb1942630 · outbound

This paper cites URL: " 'urlintro :=.

Fusion of Detected Objects in Text for Visual Question Answering URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.253125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.253125Z digest=sha256:dbd9c03b67595503f9050e82ed9efcc3cbaf114aeb7f2c6d41a1a2e066ca569e

Observation 11b9a89c-e133-4739-97ae-342d0802d52e · outbound

This paper cites write newline.

Fusion of Detected Objects in Text for Visual Question Answering write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.259261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.259261Z digest=sha256:30dc717ca5e5d20178dce998edbae1b07ad7d1b1f5962f884e91e0c456f568b8

Observation 05c165af-d2d8-4bbb-a38e-bc3847d17c62 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.265173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.265173Z digest=sha256:615ccab47268a3f43c6cb91aa10008dca17146e07880deef78a1b069ad70c34f

Observation 31dbb148-0144-4f56-9129-5dd0b8b7adb3 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Fusion of Detected Objects in Text for Visual Question Answering Lawrence Zitnick, and Devi Parikh

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.271022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.271022Z digest=sha256:2aaeb816b86839fa5bc069ca24e6f84724b125721b6a70776124995d90072e9e

Observation 5aa9e006-f9b5-49c0-ba15-64c01e057948 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:38.016914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.276113Z digest=sha256:521f2bed686117b258d72b9a79040d32829239336c3e009b0f2624ccce965344

Observation e618b184-9f10-4eaa-9d1a-a23dc8445165 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.281716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.281716Z digest=sha256:041ef91ca9725b674726a6d8f2ea09d90dc0e875fa2f50b462cf951e1702ade5

Observation 043e9063-fc3c-418c-9d8f-f8dad886f8a4 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Fusion of Detected Objects in Text for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.288311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.288311Z digest=sha256:5f8125e4370ef21a67553d37316e0c21ff2bc04112178359c63c65593e5a1ba5

Observation 1ef2e083-dea4-4620-9190-7f0ae24888b4 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.294146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.294146Z digest=sha256:f0a6d8d499faf4e8c6ae41f05cd43a4ee89c0f2671086d6689f66e0ebc7822f3

Observation 4b65b0a6-6ca3-4e30-bde0-c601f86c88cf · outbound

This paper cites End-to-End Retrieval in Continuous Space.

Fusion of Detected Objects in Text for Visual Question Answering End-to-End Retrieval in Continuous Space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.299870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.299870Z digest=sha256:67eaa2e882435011ee95e0a969d924de73be413aed05a309c77ed7aea3a6166a

Observation 0c2b84ef-4568-424e-a289-8bda653cb6c9 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.964558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.305737Z digest=sha256:2f684e3c1235d3dae36a6a1a6f7f5d94b8a7893e5085883c49ac0e6118d44b51

Observation 8aaf8743-8927-4a47-8efa-6f36688b73b1 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.313896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.313896Z digest=sha256:356186d85fcc41c1109ba1885a4f54ef4f573e3e70e92ae73949104f13cd72c0

Observation 8e3d74d8-e3a9-432c-a3e2-12040b1d9070 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.928650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.319838Z digest=sha256:deafa1f0a8dce12864110bc5c216b4f5c9e2dc2750938ea1d22ba289996b94fe

Observation 78e29a82-1070-4b8c-909b-2fb66d1a7dfd · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.906943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.325216Z digest=sha256:942bb1a0173777fc11ddd7dbdfa596fc1383106e479380853d140771ea47b8f7

Observation 6cf61d05-c208-4829-a937-bb205bd025c2 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Fusion of Detected Objects in Text for Visual Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.330350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.330350Z digest=sha256:e8ad6cb982f8af482ad11ad75cb3a8bf56aa7747f224e854af521b18937002eb

Observation 3008f3d6-c80d-4622-8b5f-0053c7caa06f · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.335390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.335390Z digest=sha256:fe8cf1c8cfb2224ad190d1ef334cd14c7f390aaba17ca3ac19cbf097c7b672a4

Observation d4fe3b53-98ea-49c4-a1ae-10ebab73d83e · outbound

This paper cites Learning Visually Grounded Sentence Representations.

Fusion of Detected Objects in Text for Visual Question Answering Learning Visually Grounded Sentence Representations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:30:37.634650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.340406Z digest=sha256:034f135af86ed95acfdabe24f1ea944463daf78acbc5c372193a858055ed47d1

Observation 1e8560b2-c313-49af-a7f9-6dd45ffdd36a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fusion of Detected Objects in Text for Visual Question Answering Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.345838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.345838Z digest=sha256:86107eec570bca232b8976539d5669ebf936cce271fa449dd7b7e3d33c2bf27b

Observation dfbe82c1-dc26-4f58-bde9-8af72792eba4 · outbound

This paper cites Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training.

Fusion of Detected Objects in Text for Visual Question Answering Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.351291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.351291Z digest=sha256:a6921e8a8c416f5a8854f88bd95ef77a8250136e55d1af18a2901bbadbfd148f

Observation cf8c80d7-1864-4e9b-9e44-140ceefe7e79 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Fusion of Detected Objects in Text for Visual Question Answering VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.356713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.356713Z digest=sha256:7f6b8ac96015bd608ad52755b69e8c2550840c19455818245e90bc320ebd3a5f

Observation 27d52ebb-3d83-4a4e-939d-52d7f94dda0c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.362170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.362170Z digest=sha256:8148ebbc2b1a86d4d7d318cfbde5824e03eb1447dccdc0a7ec22bd5376a82528

Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.367264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.367264Z digest=sha256:3e76976fc1e5a959e175ae5112db364438106bf52bfc31d365adafed98b8b4f8

Observation 78bfafc2-37a7-4236-aa37-cc6cf814bc3d · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.854627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.372574Z digest=sha256:1e466125d778d3f6657d5badb930f7897d7b9297fcf8a9ff6e7e8a3f541912cc

Observation b11a1e46-7299-46e7-ba4e-90c9131974ea · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.377545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.377545Z digest=sha256:416d147f435b6dabceea535d76a45e3ba3688c06e77387fac7c905d4ae9a4a1a

Observation 7e46b79f-1ebe-456e-9617-bf69d3b901ec · outbound

This paper cites Ororbia, Ankur Mali, Matthew A.

Fusion of Detected Objects in Text for Visual Question Answering Ororbia, Ankur Mali, Matthew A

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:30:37.814975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.383263Z digest=sha256:a056cb2666bfa6aea9da5bbfb8f84d00fbe3f2656c0f7952392b6c23d65ddfc3

Observation 15b902d3-d93d-40a1-a3a5-7cdfa23b106c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.792555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.388635Z digest=sha256:378c60ea019e5cdf14a673bdbc0291e4ec592ff5cc4d9b50992257eb38f3bf48

Observation 7ea1d988-d974-4305-81b9-403be01ce3ef · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

Fusion of Detected Objects in Text for Visual Question Answering VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.394653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.394653Z digest=sha256:e8416c9d3ed31e65731ce7f6ca7807bd28524f27d7887e2d7ed1474dbfdf437f

Observation ec81e925-860a-474d-a44e-a849cdd4ca24 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

Fusion of Detected Objects in Text for Visual Question Answering VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.401404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.401404Z digest=sha256:4e7a1830c86158dfcf69f7a42c88376333be5953bfff423d0eab68a73fe8b26c

Observation 3496b57e-4417-4ee7-8e07-f373f648c696 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.407783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.407783Z digest=sha256:11e93177587f3d77140dc6590e0ca3fb39d44d594ca542983225d0c4a3a2c238

Observation 43783d88-15ab-4c61-9ddd-cff605fb5eec · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.413255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.413255Z digest=sha256:a0a39b4e9817e7e879a5b32a83d5197dd9b0c923edb2503296e3a916e11863cb

Observation 7d6d95d9-8619-465f-ba8e-a47dab8f830c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.743792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.418907Z digest=sha256:dee3683b882e20a21ef6015779a1e14c91ac5d97ecb07980a27c022050c19085

Observation 50db3b6f-5790-4d7b-a58b-09da03951377 · outbound

This paper cites From Recognition to Cognition: Visual Commonsense Reasoning.

Fusion of Detected Objects in Text for Visual Question Answering From Recognition to Cognition: Visual Commonsense Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.425624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.425624Z digest=sha256:60861751ea0700c145a7baf029d40d4af075482cbd4d1e95821788dc23f592ca

Observation eb3b3d02-e598-43e1-8da1-5a426f0765d6 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.724844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.432399Z digest=sha256:a3ff0493f42b3303512a9b65f9b07dddd79c35a45090050235c627ec13fa4bf2

Pith citing papers

Observation 58888512-c5e6-46c3-8956-85fd6c948f8a · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Fusion of Detected Objects in Text for Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.232217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.232217Z digest=sha256:571c8894a208768b8dc65679374ba26a29842183495917f4cf80eadf8dabec05

Observation 3a1a173a-97d1-48b2-ae03-9441c812b080 · inbound

VL-BERT: Pre-training of Generic Visual-Linguistic Representations cites this paper.

VL-BERT: Pre-training of Generic Visual-Linguistic Representations Fusion of Detected Objects in Text for Visual Question Answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:18.974728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:42:18.974728Z digest=sha256:8bbb041c92a19dad47953075a1d411902aad4ef60a81a0d7b113a14c89b2f0b2

Observation 8850bb14-7819-4336-8af0-ba5204082ab3 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Fusion of Detected Objects in Text for Visual Question Answering

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.090968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:36e6f386bc1f21f8ca58989c2a69f5838623ae4acf0e2ebebc3dea4f8db2c158