Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.432399Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:1908.05054.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:30:37.432399Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-14T13:01:13.232217Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T13:03:58.089358Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4ccb0bb9-47f5-490b-9b41-9fbfb1942630 · outbound
Fusion of Detected Objects in Text for Visual Question Answering URL: " 'urlintro :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b9a89c-e133-4739-97ae-342d0802d52e · outbound
Fusion of Detected Objects in Text for Visual Question Answering write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05c165af-d2d8-4bbb-a38e-bc3847d17c62 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31dbb148-0144-4f56-9129-5dd0b8b7adb3 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Lawrence Zitnick, and Devi Parikh
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa9e006-f9b5-49c0-ba15-64c01e057948 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation e618b184-9f10-4eaa-9d1a-a23dc8445165 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 043e9063-fc3c-418c-9d8f-f8dad886f8a4 · outbound
Fusion of Detected Objects in Text for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ef2e083-dea4-4620-9190-7f0ae24888b4 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b65b0a6-6ca3-4e30-bde0-c601f86c88cf · outbound
Fusion of Detected Objects in Text for Visual Question Answering End-to-End Retrieval in Continuous Space
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c2b84ef-4568-424e-a289-8bda653cb6c9 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8aaf8743-8927-4a47-8efa-6f36688b73b1 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e3d74d8-e3a9-432c-a3e2-12040b1d9070 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 78e29a82-1070-4b8c-909b-2fb66d1a7dfd · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6cf61d05-c208-4829-a937-bb205bd025c2 · outbound
Fusion of Detected Objects in Text for Visual Question Answering GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3008f3d6-c80d-4622-8b5f-0053c7caa06f · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4fe3b53-98ea-49c4-a1ae-10ebab73d83e · outbound
Fusion of Detected Objects in Text for Visual Question Answering Learning Visually Grounded Sentence Representations
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1e8560b2-c313-49af-a7f9-6dd45ffdd36a · outbound
Fusion of Detected Objects in Text for Visual Question Answering Adam: A Method for Stochastic Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfbe82c1-dc26-4f58-bde9-8af72792eba4 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf8c80d7-1864-4e9b-9e44-140ceefe7e79 · outbound
Fusion of Detected Objects in Text for Visual Question Answering VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27d52ebb-3d83-4a4e-939d-52d7f94dda0c · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · outbound
Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78bfafc2-37a7-4236-aa37-cc6cf814bc3d · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b11a1e46-7299-46e7-ba4e-90c9131974ea · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e46b79f-1ebe-456e-9617-bf69d3b901ec · outbound
Fusion of Detected Objects in Text for Visual Question Answering Ororbia, Ankur Mali, Matthew A
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 15b902d3-d93d-40a1-a3a5-7cdfa23b106c · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7ea1d988-d974-4305-81b9-403be01ce3ef · outbound
Fusion of Detected Objects in Text for Visual Question Answering VL-BERT: Pre-training of Generic Visual-Linguistic Representations
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec81e925-860a-474d-a44e-a849cdd4ca24 · outbound
Fusion of Detected Objects in Text for Visual Question Answering VideoBERT: A Joint Model for Video and Language Representation Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3496b57e-4417-4ee7-8e07-f373f648c696 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43783d88-15ab-4c61-9ddd-cff605fb5eec · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d6d95d9-8619-465f-ba8e-a47dab8f830c · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 50db3b6f-5790-4d7b-a58b-09da03951377 · outbound
Fusion of Detected Objects in Text for Visual Question Answering From Recognition to Cognition: Visual Commonsense Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3b3d02-e598-43e1-8da1-5a426f0765d6 · outbound
Fusion of Detected Objects in Text for Visual Question Answering Unresolved cited work
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 58888512-c5e6-46c3-8956-85fd6c948f8a · inbound
Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training Fusion of Detected Objects in Text for Visual Question Answering
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a1a173a-97d1-48b2-ae03-9441c812b080 · inbound
VL-BERT: Pre-training of Generic Visual-Linguistic Representations Fusion of Detected Objects in Text for Visual Question Answering
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8850bb14-7819-4336-8af0-ba5204082ab3 · inbound
DragNUWA: Fine-grained Control in Video Generation by Integrating Text, Image, and Trajectory Fusion of Detected Objects in Text for Visual Question Answering
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.