Pith. sign in

Paper Citation Record · LEDGER

Multi-modality Latent Interaction Network for Visual Question Answering

As of 15 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 1 inbound Pith citation observation for arXiv:1908.04289.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04289 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:10:16.760619Z

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.530287Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-14T12:22:25.827597Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact3
  • verified fuzzy35
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f8794b00-5993-4e9f-9142-22eebf63d381 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Bottom-up and top-down attention for image captioning and visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.556713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.516954Z digest=sha256:0acf3880fa8b8f705a964dc1840d597f8ffebc02dd4c52f0481b8b1a25fb9cad

Observation 3941720b-7a5b-43cb-80ca-b9bd5042be37 · outbound

This paper cites Vqa: Visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Vqa: Visual question answering

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.542685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.522005Z digest=sha256:05cbd7402eba2cb5fb827381d31693ad67a59bdd15b3a07d59970427bda64482

Observation 52832828-3328-471d-bdde-897415fba561 · outbound

This paper cites Mutan: Multimodal tucker fusion for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Mutan: Multimodal tucker fusion for visual question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.527839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.526459Z digest=sha256:b6a7dd7e40049053fc44d4944b6fcd4bc80cb06d609fa7c392e912fe81e7126a

Observation d7fc90f7-b5cc-4559-9906-8d368a04006d · outbound

This paper cites Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning.

Multi-modality Latent Interaction Network for Visual Question Answering Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.531280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.531280Z digest=sha256:e98a478b8b7318981b48669aa9c7b2b1da28509009add3f73e55ba42dd7e4945

Observation 63c24c3a-8e2f-48bf-b099-1c4c1ef00846 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Multi-modality Latent Interaction Network for Visual Question Answering Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.506296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.535373Z digest=sha256:09992fa745a575c37c2e17cf1cababda9bf4c1963b9396ddab99c40be2081bef

Observation 61c023c2-b3af-4047-95df-fd6490f84d96 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Multi-modality Latent Interaction Network for Visual Question Answering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.539555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.539555Z digest=sha256:feb601bad1d81622d8f11356d620c15a2e44fd5543d990fdee47b1e4920d8d27

Observation c7df4cfc-ad11-4d94-9eea-30cb792f96f3 · outbound

This paper cites Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding.

Multi-modality Latent Interaction Network for Visual Question Answering Multimodal Compact Bilinear Pooling for Visual Question Answering and Visual Grounding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.544173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.544173Z digest=sha256:4b909ba6e26c3b8166b4d7e89425f8c19b58bddfcf66a9431127733cdf5c519b

Observation 532a131c-4239-4c17-956d-5d8bde91264d · outbound

This paper cites Dy- namic fusion with intra-and inter-modality attention flow for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Dy- namic fusion with intra-and inter-modality attention flow for visual question answering

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.492429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.549038Z digest=sha256:07f039ac88ca82062e99aaa4b79ec8def27cf84772f068fdad10ba12c542bb60

Observation f626cac9-74e9-4d54-8c40-7c3694423d9b · outbound

This paper cites Question-guided hy- brid convolution for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Question-guided hy- brid convolution for visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.478287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.553086Z digest=sha256:e85a8a46309f4cf243ed1b566df87aeac7f43ee9a564e326dac685f53a81acc6

Observation 8153b2b5-04b5-446e-a30f-0a7f0c59ab12 · outbound

This paper cites Compact bilinear pooling.

Multi-modality Latent Interaction Network for Visual Question Answering Compact bilinear pooling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.462487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.556968Z digest=sha256:b56d728c7e3cf9271ec25189b6c3e1f43e0344c954d796c644cc2e9f98ff99e1

Observation 90176aa6-a73a-4583-9813-cbf714d5130b · outbound

This paper cites 2nd place solution to the gqa challenge.

Multi-modality Latent Interaction Network for Visual Question Answering 2nd place solution to the gqa challenge

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.448325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.561528Z digest=sha256:fc506a8b335a057bc38e2656bbe316cce0042da5584c45a5701d32ed5ac1b327

Observation d74ce0eb-9090-44eb-9b36-e42a761c56a5 · outbound

This paper cites Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Making the V in VQA matter: Ele- vating the role of image understanding in Visual Question Answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.434668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.570021Z digest=sha256:60a2525cc27c8a6bfeffe69bbf85e213b6b2197ce6e6fd7ac625bed2fcb1ae43

Observation 88e27c44-0f68-48ef-8eb1-1cef736b7665 · outbound

This paper cites Deep residual learning for image recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Deep residual learning for image recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.420596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.573779Z digest=sha256:135267555795969abe2b37bccbfbb753c5a38eb082a1e8616fd5691c966693eb

Observation 6a995f6b-c0bd-433d-a828-a4f9984090a5 · outbound

This paper cites Relation networks for object detection.

Multi-modality Latent Interaction Network for Visual Question Answering Relation networks for object detection

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.405333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.577727Z digest=sha256:2ade0672e7b88f1965f586c65013396f00473900e6ab79e1d5e8dfa6871ba4ed

Observation 4fcf0b6b-17f1-4c3c-acad-41557ff3f9be · outbound

This paper cites Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Weakly-supervised Compositional FeatureAggregation for Few-shot Recognition

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.921728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.581579Z digest=sha256:b2d14baf9b1eb3ae5157210c40d0a3733288cb46ff56292058a5039edd6755ba

Observation 478c626e-c65b-4415-a993-59441475868e · outbound

This paper cites Learning to segment every thing.

Multi-modality Latent Interaction Network for Visual Question Answering Learning to segment every thing

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.392561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.585698Z digest=sha256:91a16b6496798ac948c253b2706ee631559c19469e4a78d3f81e7807c4185a81

Observation c800019c-d2df-4c7e-97fd-1769b900edfc · outbound

This paper cites Densely connected convolutional net- works.

Multi-modality Latent Interaction Network for Visual Question Answering Densely connected convolutional net- works

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.589671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.589671Z digest=sha256:d072ea11087f58fae53db2ab896f3fc5ff33d6772adb4349148a85cf48729803

Observation 7ffb9b8b-e6e0-4ff7-aa07-2c926e53558b · outbound

This paper cites Video object detection with locally-weighted deformable neighbors.

Multi-modality Latent Interaction Network for Visual Question Answering Video object detection with locally-weighted deformable neighbors

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.372598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.594324Z digest=sha256:a3fabaa874b097cb8619cc04ff28ac503387214520d707922732536f90bfa33a

Observation e35de43b-3009-4694-9a00-0665efae6991 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning.

Multi-modality Latent Interaction Network for Visual Question Answering Clevr: A diagnostic dataset for compositional language and elemen- tary visual reasoning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.359045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.598794Z digest=sha256:401eb3a1b8df929feb284338b75ea5ecd41cd3ce2464b588acb648dd0b473ee4

Observation 7f53e586-a14c-4007-91de-fa7bf28b6dcc · outbound

This paper cites An analysis of visual question answering algorithms.

Multi-modality Latent Interaction Network for Visual Question Answering An analysis of visual question answering algorithms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.340027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.603448Z digest=sha256:e43659c0150b1fc4a293f4fadbcf2cba54863e5a7a58fda96a7004aa64e70f02

Observation c2d20e82-c705-4789-8e7b-4f4a32468a76 · outbound

This paper cites Bilin- ear attention networks.

Multi-modality Latent Interaction Network for Visual Question Answering Bilin- ear attention networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.324081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.607374Z digest=sha256:5470d2ca39b1aa037f4695964b68954a4b9dd64045da5aa4b604020fb5128a91

Observation 789e25c9-c14a-47d9-9448-52a1758c08d1 · outbound

This paper cites Hadamard Product for Low-rank Bilinear Pooling.

Multi-modality Latent Interaction Network for Visual Question Answering Hadamard Product for Low-rank Bilinear Pooling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.611224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.611224Z digest=sha256:15d540780fdb8e38506597bd0ac33dfde60da0f54fedffaa1c8928bfb22b7b2d

Observation d7099b22-edbf-4701-875c-c95a3db858fa · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multi-modality Latent Interaction Network for Visual Question Answering Adam: A Method for Stochastic Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.615304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.615304Z digest=sha256:41fb8e01fdb4124891ebc84dae7b0bcbc1dbb9c3ddc969c7abd44ddef8640b97

Observation fba9b756-99a7-4e13-9aa4-b1a17e9f3783 · outbound

This paper cites Skip-thought vectors.

Multi-modality Latent Interaction Network for Visual Question Answering Skip-thought vectors

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.308746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.621107Z digest=sha256:6eb35d27e711550339663285aca119a854c9e2ba92d40f3fde73a4da5347fe5c

Observation 6cfde9e3-b90f-4126-8fe4-a12cc0d4b3ed · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Multi-modality Latent Interaction Network for Visual Question Answering Imagenet classification with deep convolutional neural net- works

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.625287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.625287Z digest=sha256:9ef6c570403f6068dcb28c0b72c2576a7c9cefc42cfe5fee892ae233c852e767

Observation 567d933e-59d2-4345-9993-f8dfe5c490fa · outbound

This paper cites Microsoft coco: Common objects in context.

Multi-modality Latent Interaction Network for Visual Question Answering Microsoft coco: Common objects in context

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.629337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.629337Z digest=sha256:91371cd95a2c730b40311a7e6b5a043e2e30c8550a2c0c43cea15fd34147d3f0

Observation 7b111dbf-8d54-49fd-9a3e-3b66434d39ef · outbound

This paper cites Improving referring expression grounding with cross-modal attention-guided erasing.

Multi-modality Latent Interaction Network for Visual Question Answering Improving referring expression grounding with cross-modal attention-guided erasing

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.264602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.633455Z digest=sha256:70e1f1b475e7c3e3f57bfc9bd277cbcbc71882bed71e78259f81c5d6de3b9068

Observation fe831b02-d01a-4554-b47e-7eee89b55662 · outbound

This paper cites ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks.

Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.637749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.637749Z digest=sha256:1535704643e5cb82c60cbf6c3f96433dcd82a3ab26e8d32dc9a21cb1729916df

Observation d1e91f63-aeee-4e88-af30-73411c56a52c · outbound

This paper cites Hierarchical question-image co-attention for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Hierarchical question-image co-attention for visual question answering

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.245332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.642418Z digest=sha256:895972dea834c93aa2fd11e678b569218a7750a1a52a7482acdbe85828887ec1

Observation 25361de9-a941-441c-b375-964800efefb6 · outbound

This paper cites Distributed representations of words and phrases and their compositionality.

Multi-modality Latent Interaction Network for Visual Question Answering Distributed representations of words and phrases and their compositionality

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.228948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.646144Z digest=sha256:8ed7f8fb6f9c55728d83570bd98e38f18428116288c39afd0ffeb6afa3bc99ea

Observation 1a56c1bf-1028-4e01-a1ed-d84b7ac9c4cd · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Improved fusion of visual and language representations by dense symmetric co-attention for visual question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.215859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.649799Z digest=sha256:7a3e4b1a9c04b23a29881b383dc3a82de9ddd7a307a187cf5b9fadf9fb26ef27

Observation cea23f11-e8c5-4b1f-a78f-7a5d3a403b19 · outbound

This paper cites Training Recurrent Answering Units with Joint Loss Minimization for VQA.

Multi-modality Latent Interaction Network for Visual Question Answering Training Recurrent Answering Units with Joint Loss Minimization for VQA

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.653933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.653933Z digest=sha256:db43dc0cb9ae52dc66476ebb47ef45caff5aad6ccfd905f85e02bf67eabf3aa4

Observation 60aacede-dd90-40e1-81e4-0b32d574e339 · outbound

This paper cites Im- age question answering using convolutional neural network with dynamic parameter prediction.

Multi-modality Latent Interaction Network for Visual Question Answering Im- age question answering using convolutional neural network with dynamic parameter prediction

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.201769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.658486Z digest=sha256:24abc7b41997e366d968180904a640b364bd248aa5af12cbd530e76c0c4148f9

Observation 988bf229-828f-48e0-8335-df53c9ccadca · outbound

This paper cites Learning conditioned graph structures for interpretable vi- sual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Learning conditioned graph structures for interpretable vi- sual question answering

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.185787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.664105Z digest=sha256:fd0e1722780bddd9f0c5cc15c32e05d58c57854aa50c9013df0fdc0b634b0604

Observation 9b0fb6fd-9f3c-4f04-9c4f-f11874844548 · outbound

This paper cites Automatic differentiation in pytorch.

Multi-modality Latent Interaction Network for Visual Question Answering Automatic differentiation in pytorch

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.668428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.668428Z digest=sha256:9bef0db92ffdbb61cdc36f767f939faf4ecc989a52e94a5004530c0c635d73af

Observation 40a778a7-310d-451a-b131-44197d020028 · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.674020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.674020Z digest=sha256:c87c7d9a1568d9a20cf978783e6b7229533b188a5343eaf34e0019a4c4670c6d

Observation ce22324b-860a-4eae-9220-d29dbf2e4eeb · outbound

This paper cites Glove: Global vectors for word representation.

Multi-modality Latent Interaction Network for Visual Question Answering Glove: Global vectors for word representation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.678574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.678574Z digest=sha256:7a29788eab0a7805b0fb4e34c01a112e4198cb02db2740336ad11a370496242f

Observation 68e5d917-76b3-4481-877f-3c6e85b6a2b3 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer.

Multi-modality Latent Interaction Network for Visual Question Answering Film: Visual reasoning with a general conditioning layer

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.146576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.683012Z digest=sha256:44742bc2d9cff3273aba3e549f17e62267464101f149c9ce3846aec10f4cdd86

Observation 7522a3e6-cf18-4dde-80a8-d9cf8eb13b5c · outbound

This paper cites Deep contextualized word representations.

Multi-modality Latent Interaction Network for Visual Question Answering Deep contextualized word representations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.130481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.688219Z digest=sha256:14dfd15c7a236d1a82deb61d5ae435ebb7b1f2b2fb90ae77d06da13b49760e28

Observation b48ce64b-0f92-4ee7-a237-1fec3d7f0918 · outbound

This paper cites Language models are unsuper- vised multitask learners.

Multi-modality Latent Interaction Network for Visual Question Answering Language models are unsuper- vised multitask learners

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.692191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.692191Z digest=sha256:8b2a46a8a5826f4f9d6fbc60a81519b978cd8a3c672807cffc4a063be52b7cd8

Observation 52362887-7d93-44d1-b0ee-0963074959e1 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

Multi-modality Latent Interaction Network for Visual Question Answering Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.109313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.696716Z digest=sha256:74d071d5962c49bef04a421705a0aecd4a4936f853951fe642385a75dc65d561

Observation 73f4e08b-e4aa-419a-9b1e-9953aa6767ac · outbound

This paper cites A simple neural network module for relational rea- soning.

Multi-modality Latent Interaction Network for Visual Question Answering A simple neural network module for relational rea- soning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.094891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.701080Z digest=sha256:d135a0f63b80975f703f9b7fcd8bb97d4361028916f2709f999337934c4af21e

Observation 2d6bca56-b516-466c-b929-64a53a53c5bf · outbound

This paper cites Question type guided attention in visual ques- tion answering.

Multi-modality Latent Interaction Network for Visual Question Answering Question type guided attention in visual ques- tion answering

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.081647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.705116Z digest=sha256:32dd7904c72680acf89233ea4eec830deb68a737b1ab3d12fe6446609112b9e0

Observation 85f2cdf2-0d6e-4cbc-97fe-76967bd6ec6a · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Multi-modality Latent Interaction Network for Visual Question Answering Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.709992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.709992Z digest=sha256:bb03aa224768b8af79d0bf76f4c2d0bad5f6a1bd8c97356adbd46a8a1a28dfb1

Observation b6ee223f-13c4-4ed4-b049-003d9f6bf967 · outbound

This paper cites Attention is all you need.

Multi-modality Latent Interaction Network for Visual Question Answering Attention is all you need

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.069512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.713848Z digest=sha256:8ab2abe270102adca6bb14e7a498c3c0f3838fb1e5838c68c0e7ad885bbca96a

Observation f9aa4ee0-2de4-4ff1-9ddb-b916780cc722 · outbound

This paper cites Non-local neural networks.

Multi-modality Latent Interaction Network for Visual Question Answering Non-local neural networks

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.717747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.717747Z digest=sha256:1fc200c9c9787b969b45bce75333e3753d9ec0f7e8be0d7a8d5439f333e408d0

Observation e530c4c5-f0de-477b-8477-f737887aa4dd · outbound

This paper cites Pay Less Attention with Lightweight and Dynamic Convolutions.

Multi-modality Latent Interaction Network for Visual Question Answering Pay Less Attention with Lightweight and Dynamic Convolutions

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.722235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.722235Z digest=sha256:1ac8bb9bae8162b297dc88f51d7a1948c5e1873b160d5c04f88caedc4d19e2ac

Observation c1ff441c-aa0b-41ce-b19d-037395704ed9 · outbound

This paper cites Show, attend and tell: Neural image caption gen- eration with visual attention.

Multi-modality Latent Interaction Network for Visual Question Answering Show, attend and tell: Neural image caption gen- eration with visual attention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.047972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.726466Z digest=sha256:6c0a6ae776d8c55fab03640403be748d6e54cd97fbf3484a00fe9fd3ee03e88e

Observation 00a67c6a-a086-41c4-ad6c-cbc161570100 · outbound

This paper cites Stacked attention networks for image question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Stacked attention networks for image question answering

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.034059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.731155Z digest=sha256:2c10ff23b7a4ca1063363c47a6856b7f225b9ad61c6f33fa9a19ff4261ecf9c1

Observation e82286dc-ef8a-4aab-aedb-93690a8ffbb2 · outbound

This paper cites Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Scene Graph Reasoning with Prior Visual Relationship for Visual Question Answering

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.814204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.735502Z digest=sha256:9cfbd3020d2a05e48f668de478c9f2a4d702d72d36d99364e7c89a7d794169fb

Observation 7ec17cdb-760a-4643-b1c0-c87333ed883f · outbound

This paper cites Explor- ing visual relationship for image captioning.

Multi-modality Latent Interaction Network for Visual Question Answering Explor- ing visual relationship for image captioning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.018843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.739200Z digest=sha256:aa0507d7f3ee9467298713b73c84964f14e9d763956881d98c77913d3ee180b8

Observation feb6f77b-7d61-4ecb-b53b-76d295491612 · outbound

This paper cites Beyond bilinear: generalized multimodal factorized high-order pooling for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Beyond bilinear: generalized multimodal factorized high-order pooling for visual question answering

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:17.006135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.743986Z digest=sha256:426aebe20916e6778c30938391ae495fe676ea122400fdc4d3c26e9ad9bc48fa

Observation de7e3c54-6ffb-4a8d-a514-38cbd2ab6e8f · outbound

This paper cites Yin and Yang: Balancing and an- swering binary visual questions.

Multi-modality Latent Interaction Network for Visual Question Answering Yin and Yang: Balancing and an- swering binary visual questions

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:16.991110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.747846Z digest=sha256:11c720a8c65f869e8f5709743411cfa763dac4375dd45e6bfd83aa2b40765ea6

Observation f4fcaa43-606e-477c-a945-f8fe8977609d · outbound

This paper cites Learning to Count Objects in Natural Images for Visual Question Answering.

Multi-modality Latent Interaction Network for Visual Question Answering Learning to Count Objects in Natural Images for Visual Question Answering

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.755399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.755399Z digest=sha256:425bdc37aed81b255d45939fa3d7ed0233b50d7495d56366e4a9b6ae840ad50c

Observation 57c5b259-ddc5-4643-99f5-a6f8b52f7753 · outbound

This paper cites Structured attentions for visual question answering.

Multi-modality Latent Interaction Network for Visual Question Answering Structured attentions for visual question answering

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:10:16.976330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.760619Z digest=sha256:9effb4f27eb038a48204f814ff4475ccbcb5e4e4bbfbcb1d4367248d152b0e10

Observation abf6c8a0-e384-41c4-a8c4-1eeea82551cd · outbound

This paper cites 2nd Place Solution to the GQA Challenge 2019.

Multi-modality Latent Interaction Network for Visual Question Answering 2nd Place Solution to the GQA Challenge 2019

Reference 2019

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:10:16.939521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T14:10:16.565726Z digest=sha256:cc627052d7ff6ff0380e58d7a8232b4480bb01f1acd9786683c0fcfbe800b3cb

Pith citing papers

Observation 32e70082-8c07-4327-8b3d-e21b107212c7 · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multi-modality Latent Interaction Network for Visual Question Answering

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:22:25.833918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-14T12:22:25.530287Z digest=sha256:ace0c3911d6fc176ec156870b69ab6e06b2bbd93638719f066b3c114cd8e08c0