Pith. sign in

Paper Citation Record · LEDGER

Fusion of Detected Objects in Text for Visual Question Answering

As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:1908.05054.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.05054 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.432399Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:01:13.232217Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T13:03:58.089358Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4ccb0bb9-47f5-490b-9b41-9fbfb1942630 · outbound

This paper cites URL: " 'urlintro :=.

Fusion of Detected Objects in Text for Visual Question Answering URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.253125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.253125Z digest=sha256:ac92f121b87f5638ee0e36d34f70b666d78a611bfa9322e418ad1b8999f29a06

Observation 11b9a89c-e133-4739-97ae-342d0802d52e · outbound

This paper cites write newline.

Fusion of Detected Objects in Text for Visual Question Answering write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.259261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.259261Z digest=sha256:16dfedef4abaefb0f252d163a1799d8667e68b6e35b8cd64b95618d195d64b7f

Observation 05c165af-d2d8-4bbb-a38e-bc3847d17c62 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.265173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.265173Z digest=sha256:b8701d857f00d8662ad6c132eb8d532ee3feae409945217921bf3ebfab558300

Observation 31dbb148-0144-4f56-9129-5dd0b8b7adb3 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

Fusion of Detected Objects in Text for Visual Question Answering Lawrence Zitnick, and Devi Parikh

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.271022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.271022Z digest=sha256:1882f119d2f4f263d5e95e3e7129bf3d489ad22f15beb4cb57c1ffe89c422575

Observation 5aa9e006-f9b5-49c0-ba15-64c01e057948 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:38.016914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.276113Z digest=sha256:2d248eaced6274efd3460ac3bf2b8dadf17b0699da50631926b73d0bb59091b5

Observation e618b184-9f10-4eaa-9d1a-a23dc8445165 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.281716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.281716Z digest=sha256:153943d08f92d22c07b478fcdb0b0b6cab3acccb65035358c6c6a7d9bc5a13db

Observation 043e9063-fc3c-418c-9d8f-f8dad886f8a4 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Fusion of Detected Objects in Text for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.288311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.288311Z digest=sha256:07a8763ef8ec8e3de3d6dd188aacc9c2e46a2cd92c22de6d0da6e10ceebda721

Observation 1ef2e083-dea4-4620-9190-7f0ae24888b4 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.294146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.294146Z digest=sha256:88eb5d3eb5dc06caa149c2287a5ad8bb66f2f9188e52c8d6a2fa51434d65fe35

Observation 4b65b0a6-6ca3-4e30-bde0-c601f86c88cf · outbound

This paper cites End-to-End Retrieval in Continuous Space.

Fusion of Detected Objects in Text for Visual Question Answering End-to-End Retrieval in Continuous Space

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.299870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.299870Z digest=sha256:a89f8f432229f5685954aaa12fe925bdd285b7398aeb5dfa77d811b1a3f0c945

Observation 0c2b84ef-4568-424e-a289-8bda653cb6c9 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.964558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.305737Z digest=sha256:da754a0c3430e4604a7a9223bd0a17bd885de12d980bc7cc48d2c35b973ae2ab

Observation 8aaf8743-8927-4a47-8efa-6f36688b73b1 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.313896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.313896Z digest=sha256:f26fcd41a897a4ec80276741d9d7946f37d0a51ce27b6a302dc2ba2f02abf7e5

Observation 8e3d74d8-e3a9-432c-a3e2-12040b1d9070 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.928650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.319838Z digest=sha256:78b568854d9cf17a5e1b71e58989b5dfa7cda77dc43758dda9f747d02bdf667b

Observation 78e29a82-1070-4b8c-909b-2fb66d1a7dfd · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.906943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.325216Z digest=sha256:3633526090746bfe66e6978a017847b724a17a0c973ac6962179d219e2c37e96

Observation 6cf61d05-c208-4829-a937-bb205bd025c2 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

Fusion of Detected Objects in Text for Visual Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.330350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.330350Z digest=sha256:d7865a69f52f0e8c8aca139cb38829fe8e380621bf90745a09cc15dc20a792bd

Observation 3008f3d6-c80d-4622-8b5f-0053c7caa06f · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.335390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.335390Z digest=sha256:9fd3047ef1635b554a8f565f1c6bfbaecb067f74c61011da33ff90574b1d2334

Observation d4fe3b53-98ea-49c4-a1ae-10ebab73d83e · outbound

This paper cites Learning Visually Grounded Sentence Representations.

Fusion of Detected Objects in Text for Visual Question Answering Learning Visually Grounded Sentence Representations

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:30:37.634650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.340406Z digest=sha256:649c99944b7a7b050dadaf928d0acc21ea11d86d26031eb1004b9f12e65ae183

Observation 1e8560b2-c313-49af-a7f9-6dd45ffdd36a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Fusion of Detected Objects in Text for Visual Question Answering Adam: A Method for Stochastic Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.345838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.345838Z digest=sha256:62958dbc060a54ff0cb2848e9cbeb092ae768ecfcc40e6fc1dfd50abb1847c0f

Observation dfbe82c1-dc26-4f58-bde9-8af72792eba4 · outbound

This paper cites Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training.

Fusion of Detected Objects in Text for Visual Question Answering Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.351291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.351291Z digest=sha256:589ebc5a41c23d6ebbecc44f8cc50fd65a624e3d18f68ce21b20bb0fc043ed7f

Observation cf8c80d7-1864-4e9b-9e44-140ceefe7e79 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

Fusion of Detected Objects in Text for Visual Question Answering VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.356713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.356713Z digest=sha256:aa5a22b2acecddd590344cea5044e784be8c757316ac0c76eb726730aeecea7c

Observation 27d52ebb-3d83-4a4e-939d-52d7f94dda0c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.362170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.362170Z digest=sha256:4591f975fe488c906acbb8041e76cc5bfd47ec753133f87e456dccf112c8e6b6

Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.367264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.367264Z digest=sha256:44a1b4cef44667fdd7d78530d55b28fa4683c0c3f5a2f095da83e347c2a82736

Observation 78bfafc2-37a7-4236-aa37-cc6cf814bc3d · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.854627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.372574Z digest=sha256:62e1185695ec87e41111342330ed4e73366f5cba7c2cd28c0c8206bc1a9e4140

Observation b11a1e46-7299-46e7-ba4e-90c9131974ea · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.377545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.377545Z digest=sha256:4e50e9949d9e1bb1eeb5ee8fe050f1c1e03a1b671e9e183f7450ab5916236d67

Observation 7e46b79f-1ebe-456e-9617-bf69d3b901ec · outbound

This paper cites Ororbia, Ankur Mali, Matthew A.

Fusion of Detected Objects in Text for Visual Question Answering Ororbia, Ankur Mali, Matthew A

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:30:37.814975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.383263Z digest=sha256:187a0cdc67d009acc45c629b9cc427e6282de770545183b01ba244620536cf3d

Observation 15b902d3-d93d-40a1-a3a5-7cdfa23b106c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.792555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.388635Z digest=sha256:2fc503daaace56c2fd62440445a7e766c5af0e552babde44467bd437a62ad30d

Observation 7ea1d988-d974-4305-81b9-403be01ce3ef · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

Fusion of Detected Objects in Text for Visual Question Answering VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.394653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.394653Z digest=sha256:75d4a9e8dc3a0f21a6695f30c52f99d4bf4eb7abb591d2e0a4f6640904d3d11d

Observation ec81e925-860a-474d-a44e-a849cdd4ca24 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

Fusion of Detected Objects in Text for Visual Question Answering VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.401404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.401404Z digest=sha256:b8d89387b2c2d5560db4a66d11208cb1baf26d420e1aaa816b65eeed94b56574

Observation 3496b57e-4417-4ee7-8e07-f373f648c696 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.407783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.407783Z digest=sha256:a796c42b234975ba092051a0f2df5ed34508ed60f52704ab50d4d2dfc56a59ca

Observation 43783d88-15ab-4c61-9ddd-cff605fb5eec · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.413255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.413255Z digest=sha256:1ef8f76cd905461bf563ff4d76bc56715d65a298fafa3f34f8017f19fb51dbe9

Observation 7d6d95d9-8619-465f-ba8e-a47dab8f830c · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.743792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.418907Z digest=sha256:46450b50519e0b6267487c41a904d7ed1689956ce9715476908094b90e3e91d2

Observation 50db3b6f-5790-4d7b-a58b-09da03951377 · outbound

This paper cites From Recognition to Cognition: Visual Commonsense Reasoning.

Fusion of Detected Objects in Text for Visual Question Answering From Recognition to Cognition: Visual Commonsense Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.425624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.425624Z digest=sha256:d7c2a5830ec80f23d47a3dd67eea76dd7bc06062aeb42dc7e67a4dcde4b512fb

Observation eb3b3d02-e598-43e1-8da1-5a426f0765d6 · outbound

This paper cites an unresolved cited work.

Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-14T13:30:37.724844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T13:30:37.432399Z digest=sha256:9828934e21613d871d443ecf1ca3e8882cb1f3455657582b0e3e49a6cc1ed2f7

Pith citing papers

Observation 58888512-c5e6-46c3-8956-85fd6c948f8a · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Fusion of Detected Objects in Text for Visual Question Answering

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.232217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.232217Z digest=sha256:5761918d2a8a81fe2fc4d06347251a5e5765c0a0ed2b80175aabbbf52521901d

Observation 3a1a173a-97d1-48b2-ae03-9441c812b080 · inbound

VL-BERT: Pre-training of Generic Visual-Linguistic Representations cites this paper.

VL-BERT: Pre-training of Generic Visual-Linguistic Representations Fusion of Detected Objects in Text for Visual Question Answering

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:18.974728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:42:18.974728Z digest=sha256:ad2c35ab7ff5faf605dfbd645460e125e2ae0f53edb5d5e0f1dc532a3f192295

Observation 8850bb14-7819-4336-8af0-ba5204082ab3 · inbound

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory cites this paper.

DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Fusion of Detected Objects in Text for Visual Question Answering

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-20T13:03:58.090968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-20T13:03:57.828598Z digest=sha256:cefc080683e3fc6f3ad1426ab04a1dc4fe0b9e5cddbbed854cac2a7f7e4b11de