Pith. sign in

Paper Citation Record · LEDGER

Multimodal Unified Attention Networks for Vision-and-Language Interactions

As of 16 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 3 inbound Pith citation observations for arXiv:1908.04107.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.04107 v2

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T13:55:26.757084Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:22:25.662003Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:39:30.840356Z

Reference resolution

70 of 70 outbound references displayed

  • verified exact1
  • verified fuzzy53
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9783f87-858d-4eb4-812b-65d09ad8df0b · outbound

This paper cites Multimodal deep network embedding with integrated structure and attribute information,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal deep network embedding with integrated structure and attribute information,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:28.008716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.405870Z digest=sha256:37839150f2c8eeb8e3614da35098bcd7617b23e05614c4f23f0e8ff552aa8c5e

Observation 70e901c3-cc0d-40bc-a28c-5f92950e3cc7 · outbound

This paper cites Discrim- inative coupled dictionary hashing for fast cross-media retrieval,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Discrim- inative coupled dictionary hashing for fast cross-media retrieval,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.991927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.411450Z digest=sha256:bc0b294bbe90be080545bfadf9a8c21a10976c31dd2a72242b437795dbc79283

Observation f6859418-36ed-429d-b6ab-d99ffa1366d7 · outbound

This paper cites Shared predictive cross-modal deep quantization,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Shared predictive cross-modal deep quantization,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.974963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.416382Z digest=sha256:2bc397cbe09bd7b34014682fea6bada25346d1ccb21d07b3df428b59deaa39da

Observation 8b5b2dbe-a68f-456a-b955-3ba975161532 · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Show, attend and tell: Neural image caption generation with visual attention

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.957946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.422009Z digest=sha256:4371de164beff97e96eee92d5f31f0503e6901bb8b43e2901d6430df2b7f0d5c

Observation 3c17dd8c-eeeb-4498-98a4-fcdeaf39fe57 · outbound

This paper cites From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions From deterministic to generative: multi-modal stochastic rnns for video cap- tioning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.940902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.427042Z digest=sha256:194138ff1e0582ef5eadf7edad8b4cce4f17e3aa31bef44a1dd1f5a850e6e94b

Observation 24dfb2cb-350e-48d1-8b1c-330066e7bc33 · outbound

This paper cites Vqa: Visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Vqa: Visual question answering,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.923633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.432661Z digest=sha256:385f7ffaff2be53ea37fffd655a8a9a19c12b4ba4f00f94a8f1ee80c95c7f5a1

Observation a712be8a-2deb-440c-8c4c-31ffb95e8ab3 · outbound

This paper cites Ground- ing of textual phrases in images by reconstruction,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ground- ing of textual phrases in images by reconstruction,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.907818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.439241Z digest=sha256:17375f1bccf8d3222949f34cd9ab9496c69402438e29d541fa65e9c1b1c6efd9

Observation 8391a958-710e-4be9-8f55-50c76d8493db · outbound

This paper cites Neural Machine Translation by Jointly Learning to Align and Translate.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Neural Machine Translation by Jointly Learning to Align and Translate

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.444239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.444239Z digest=sha256:2465ce23426af7e53a6082e93e5a28c48b423ac107b3d89892e07739859f5312

Observation d9c99a69-6c26-4d21-9df4-79d3556906e4 · outbound

This paper cites Recurrent models of visual attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Recurrent models of visual attention,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.891612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.449809Z digest=sha256:f097febd5badce447be13dbb38169b13929be70f750711e926a5148a3eee1e44

Observation 76a48858-2880-49ab-904d-fddcbfed84a8 · outbound

This paper cites Draw: A recurrent neural network for image generation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Draw: A recurrent neural network for image generation,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.875785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.454446Z digest=sha256:c4916237cf19351d123e45a72c991a4955a00832ce9a3e3fd27ea556baaee282

Observation 025187f2-3715-4aa0-b7cb-2d689b66bb17 · outbound

This paper cites Attention to scale: Scale-aware semantic image segmentation,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention to scale: Scale-aware semantic image segmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.859785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.459270Z digest=sha256:5df742b79d6c23c7ff2bfe4b2959d3c0590bbf66df603c57e920fc487e81e113

Observation cc9ad11a-415f-4503-a2dd-4b836808f174 · outbound

This paper cites Effective Approaches to Attention-based Neural Machine Translation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Effective Approaches to Attention-based Neural Machine Translation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.465116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.465116Z digest=sha256:d1d274c89fbc5e42fcbed20ff7561a7a9fd04cfb1b6748bc17448d8a25d63304

Observation 7b80ff55-9a49-41b6-9d8f-8b089392bd47 · outbound

This paper cites Deep biaffine attention for neural dependency parsing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep biaffine attention for neural dependency parsing,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.842087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.470402Z digest=sha256:76caa6ee30745495849926af3044888a71e93f617a32815e19956dd312fb2fbd

Observation 83ed30e9-de93-4058-95d8-f2258e795eb8 · outbound

This paper cites A Neural Attention Model for Abstractive Sentence Summarization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Neural Attention Model for Abstractive Sentence Summarization

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.475737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.475737Z digest=sha256:23e6de7b566a433e9f5b8ab2fe9f6a673c58e07b2c34301eaaef671e2fb71bc7

Observation 5ead648f-1b4b-4cdd-b058-1cc5db435c55 · outbound

This paper cites Stacked attention net- works for image question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Stacked attention net- works for image question answering,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.824687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.481939Z digest=sha256:527f04eed28655c9928c2ce930309917102f1bc57d9a4ff800a70348ae28b1dd

Observation 24d92376-f32e-44b1-9559-3cda4b74118f · outbound

This paper cites Multimodal compact bilinear pooling for visual question answering and visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multimodal compact bilinear pooling for visual question answering and visual grounding,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.807046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.487055Z digest=sha256:b690bf813c1f22dee84f5b6c527468b129eba92fb8c82edb9ff4efb857021405

Observation 93828af5-dd98-46b0-99cb-bde6d74cd5c6 · outbound

This paper cites Hierarchical question-image co-attention for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hierarchical question-image co-attention for visual question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.790038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.491852Z digest=sha256:2f2ed0d26d0f23547f8e9de441e1729358866e92dd89e6a934b3d41bdd383ec8

Observation 79de54b6-0d38-4847-8c0d-9db88de40266 · outbound

This paper cites Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Multi-modal factorized bilinear pooling with co-attention learning for visual question answering,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.771611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.496532Z digest=sha256:5e1a7ff464a28c20787f917cca900dffafb26031d1acc91faf03dc942e1f688c

Observation 2217e021-5f31-4498-b185-6e8a8e8ec933 · outbound

This paper cites Bilinear attention networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bilinear attention networks,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.754499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.501649Z digest=sha256:40a2d236487b0dc9d8c231f470088e00e6975da9db5ff0c7ca352d4e4bc78184

Observation b0fe8842-fcf6-4eea-ba58-794f04ddea9e · outbound

This paper cites Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improved fusion of visual and language representations by dense symmetric co-attention for visual question an- swering,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.731604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.506742Z digest=sha256:b95f91fc0b1fdc8408928aba8ad54bdbcbe4eecac025263f9e47dcb4ed14960f

Observation 8d40dc6e-7186-4ff5-bb03-b5adeb46d1d9 · outbound

This paper cites Attention is all you need,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Attention is all you need,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.710634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.513064Z digest=sha256:1839b453cf411072f469a3b64725978fc43c9cd38e10f51d009973bc374eeeb3

Observation e4dea28e-54d3-44f4-9553-ddbda29ecfb3 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Multimodal Unified Attention Networks for Vision-and-Language Interactions BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.518164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.518164Z digest=sha256:5ea3025ef2e4c647a7bff7ae340b025cefb37fd47cfe229118a08a97eae98e3e

Observation c50333a2-4c3f-4f51-8dc1-1ac7acbcc3b1 · outbound

This paper cites Non-local neural net- works,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Non-local neural net- works,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.692996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.523097Z digest=sha256:9dfc9c715970d7b538b8cef0d90506c36a38a25c990af1b9c06a3d72d5f0de3f

Observation 4d6962c0-7f30-4f83-80b5-ac0d6d0eff6b · outbound

This paper cites Relation networks for object detection,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Relation networks for object detection,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.674066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.527825Z digest=sha256:0185c09f45930d293b9da8abe6889573167298b91caf772a91845eda968d2cd6

Observation c4fdcde9-6287-4302-8336-8c7a46ddfa72 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Making the v in vqa matter: Elevating the role of image understanding in visual question answering,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.655703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.532554Z digest=sha256:fd5237e07bea411ff57b0af0d50516ace062acdf8f6beeb2290d2e9723c2a39d

Observation 51ff0f0b-5de1-4252-830d-3b16362c1344 · outbound

This paper cites Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Clevr: A diagnostic dataset for compositional language and elementary visual reasoning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.638537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.538915Z digest=sha256:3147595d5c57b12a0d1a054d651112bcecd6d0305548afd2eed994b6b60a87e7

Observation 6e577d9d-4dcb-4137-b9d6-1b817db4de5c · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referitgame: Referring to objects in photographs of natural scenes,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.543907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.543907Z digest=sha256:451b7853a14b84f72fae9462bd71aa90adad864f5b741ee38d631c70640c122d

Observation f30f20a7-9f63-40b1-b93a-9a8f24d590bd · outbound

This paper cites Generation and comprehension of unambiguous object descriptions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Generation and comprehension of unambiguous object descriptions,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.610143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.549092Z digest=sha256:00fdde60ef5dae35b7d2a246ab12e33c00e5b4f430e8a28743eb18cfbe5b4a86

Observation 6f797fb1-9b33-429f-a636-b0831eafd354 · outbound

This paper cites Simple Baseline for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Simple Baseline for Visual Question Answering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.554516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.554516Z digest=sha256:78252ca932620a50d6cb49d63ba558bce48131fea76ebdd076d417d3a08d2b3d

Observation 3d147a9b-af1f-4fb7-82c0-5b22177aadf3 · outbound

This paper cites Hadamard Product for Low-rank Bilinear Pooling,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Hadamard Product for Low-rank Bilinear Pooling,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.592926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.559720Z digest=sha256:2c3ffe3498b56ac7ef2c5a3f01dc7e7e82df9df7bd2d597358f679c00f8890af

Observation 6d9aaaf4-c355-4c10-ab2c-6d3ebe5326e1 · outbound

This paper cites Mutan: Multi- modal tucker fusion for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mutan: Multi- modal tucker fusion for visual question answering,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.573058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.564802Z digest=sha256:6a15f197a59252cec34bee651207b79dfc0ce3efd9a254bc77d775f6f9ba302b

Observation 6d91550f-15d5-4026-88b0-8a0b3148f76a · outbound

This paper cites ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions ABC-CNN: An Attention Based Convolutional Neural Network for Visual Question Answering

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.570303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.570303Z digest=sha256:80487fbd53d6107c36c0eb742e350c9611d1acd2345fb95ca497f9a49424ec37

Observation 109b5f20-38ee-426c-8fee-c9af7d2abde6 · outbound

This paper cites A Focused Dynamic Attention Model for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A Focused Dynamic Attention Model for Visual Question Answering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.575466Z digest=sha256:5dd530e9706f6767e814ae1e92ed42719677f03886f399325471c1466926527b

Observation bb9f24d8-e8e2-4819-8e40-6397ae70a6fa · outbound

This paper cites Where to look: Focus regions for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Where to look: Focus regions for visual question answering,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.553523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.580811Z digest=sha256:1264acbe8cdc29246a3b35299cd6105d4331927a5c1126dbb13a7358e1628468

Observation fb59863f-722b-4c6c-be73-b56a7b16b79a · outbound

This paper cites Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Beyond bilinear: Generalized multi-modal factorized high-order pooling for visual question answer- ing,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.533495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.586427Z digest=sha256:92ba661003a507cfe0fa98fe1917a7765351a20f35e6c00be5eb7e4ddb079882

Observation 960371a1-dd9e-41eb-aca8-324d0c3487c9 · outbound

This paper cites A joint speaker-listener- reinforcer model for referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A joint speaker-listener- reinforcer model for referring expressions,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.501146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.591097Z digest=sha256:00fd43c4210334698fdb3839d2cd772ad21d607b926164375d33ccb5d82e087f

Observation fd8e2279-07a3-4bb7-a782-0b2172e20416 · outbound

This paper cites Edge boxes: Locating object proposals from edges,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Edge boxes: Locating object proposals from edges,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.481300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.595646Z digest=sha256:575eda009b94f19ce479e6859217b2707a33e3b0e3ac5f0beea8e8f36c80eb48

Observation eaaee506-d5f9-455d-a353-4f846e9de4db · outbound

This paper cites Deep residual learning for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep residual learning for image recognition,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.462346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.600632Z digest=sha256:a22b7f1b9efa81bf78411600230291dfd40e14b20ac4b1b77f5c1d3d0af8ecf8

Observation 848abf49-fe14-48b4-8ed8-c3d6ab67e6a9 · outbound

This paper cites Rethinking diversified and discriminative proposal generation for visual grounding,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Rethinking diversified and discriminative proposal generation for visual grounding,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.444232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.606994Z digest=sha256:c1bbb5619c2e77b367d0e8e48de38081e3623c073624b45100b42bde54f24364

Observation 09816d98-c73d-422d-96a0-3633ac698948 · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mattnet: Modular attention network for referring expression comprehension,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.426916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.611863Z digest=sha256:0d8cdb9404eac64bfe4b52d49d752f1fd16bdaf7adafcc1fb513fdeff142b832

Observation 23a1ab66-d516-47ae-a85e-05211370ff20 · outbound

This paper cites Parallel attention: A unified framework for visual object discovery through dialogs and queries,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Parallel attention: A unified framework for visual object discovery through dialogs and queries,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.407914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.616740Z digest=sha256:0511b1911ca145c1def0dc11cf298811f12e0513b9aa5c1d738173d9da5eb491

Observation 9aee77a4-0b7a-4c8d-8d49-bcb2a792117e · outbound

This paper cites Visual grounding via accumulated attention,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual grounding via accumulated attention,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.390683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.622160Z digest=sha256:852433d98e80ac57c0169a01b031c1f8084268d87ede8db0b8f18bf7c5483aa6

Observation 619e07f7-7290-4b83-a249-6b0e2840a833 · outbound

This paper cites Be- yond rnns: Positional self-attention with co-attention for video question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Be- yond rnns: Positional self-attention with co-attention for video question answering,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.374092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.627435Z digest=sha256:5672f95b658a577ad336ab9936614542b672e1ab7396716bb4933a7bfcd42cb9

Observation b487a113-c177-42c6-920b-73ddb6fbff3b · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-08-14T13:55:26.904087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.632542Z digest=sha256:173e520bce6748ba5e7dc0a8416f8dfb18eeca0f61f584f7ec197e5c8816e069

Observation 8e1ba17d-fc08-41e7-b6a5-609798755313 · outbound

This paper cites Improving language understanding by generative pre-training,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Improving language understanding by generative pre-training,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.357158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.638333Z digest=sha256:2f530e2d7afde45a71a709bad62794c0d22d9b35e502dc95361a7b680a050279

Observation 56e1d26b-e8d1-4ce0-84c8-e33d878dfbfc · outbound

This paper cites Factorized bilinear models for image recognition,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Factorized bilinear models for image recognition,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.340667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.643030Z digest=sha256:20049268ac523c21dab856a48f51d34c39c5b15921ed31153236e4d044eaaf43

Observation 7c91eab1-0e15-4cd8-a276-5bb7285988bc · outbound

This paper cites Layer Normalization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Layer Normalization

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.647648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.647648Z digest=sha256:417e9ffc9d1ffd2bc88b70d31f93cd146a18f594b05f26daca5bb995565ea543

Observation 9c91039a-5194-42a8-a217-dc1fad636ef9 · outbound

This paper cites Glove: Global vectors for word representation.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Glove: Global vectors for word representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.322774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.652624Z digest=sha256:e1a21119e455f16f17f257c2503c979a2125d4f314d0d328f3817f50d525662d

Observation ca5f6cfb-f77c-4d5d-aa6c-22c2b451f2d7 · outbound

This paper cites Long short-term memory,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Long short-term memory,

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.657304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.657304Z digest=sha256:e1cff225aa4beb1d9498b3450ed5f7abfb2003dbb6521da9a6761ff26dfc6c89

Observation e0c6bc30-b4b8-4e0b-a17d-62af13d28368 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Bottom-up and top-down attention for image captioning and visual question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.290618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.662045Z digest=sha256:ef44a3c1c936b8e5f4a75742083f599d786b9ce8c3a355a8c89adceb56a4ca7c

Observation a01fe0f2-a863-4ebd-9edd-2c4db59909f0 · outbound

This paper cites Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Tips and Tricks for Visual Question Answering: Learnings from the 2017 Challenge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.666553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.666553Z digest=sha256:7b2aacaf212fc40587814aa6455b39cf76bd707c34e09387db20e37a8b8e4e78

Observation 9f73072f-c372-4c77-9609-1f8bc3152ae5 · outbound

This paper cites Fast r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Fast r-cnn,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.271577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.671520Z digest=sha256:dc5d6b0e9dc7fa3e8bcecca30af14a8056678e8f314c8485c35818f3bc46abe9

Observation 59e2ed05-d7a5-4aba-9e46-0538c12b7244 · outbound

This paper cites Microsoft coco: Common objects in context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Microsoft coco: Common objects in context,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.251132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.676051Z digest=sha256:9d4001cdd75b213f168e9180305d8036c4d1656941cedb62a0d567df031e5bcf

Observation 5617209c-ca98-4c12-938a-ea8847087550 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Adam: A Method for Stochastic Optimization

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.680958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.680958Z digest=sha256:c225f60499919ea5717a424b0ac33f5ec6d383efc280ecb0f293560b06f1f4ce

Observation ccee27c2-86b0-4da9-b579-2767d2a9f8b1 · outbound

This paper cites Faster r-cnn: Towards real- time object detection with region proposal networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Faster r-cnn: Towards real- time object detection with region proposal networks,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.233090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.685644Z digest=sha256:3b1a76c5656558ca34739e5f20d106a0f0e1e3223d77be445fc5c58f79018bc4

Observation 75b9dd9d-9b4e-436b-a1c0-6d5dae966155 · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.690283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.690283Z digest=sha256:fb6993560cb9a512da9f44be65d34665c0aff12280d16d8809f67f5ae85af30f

Observation a5d1f2f3-65f2-427c-a4ef-e1246434acb0 · outbound

This paper cites Compositional Attention Networks for Machine Reasoning.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Compositional Attention Networks for Machine Reasoning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.694870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.694870Z digest=sha256:ac23bbf2e9b30aa0aa11e60606c37edcaa641a260af41a34e9951566a88af7bc

Observation 4b37c205-f5f0-4f0b-87de-189dacaa1a74 · outbound

This paper cites Mask r-cnn,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Mask r-cnn,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.214156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.699662Z digest=sha256:a54b12a06d53d14c39cca80e902c263c38ca6f890ef875665801d164a2b3b16e

Observation dc40e05e-a0b8-42f8-8ecb-56e048881c88 · outbound

This paper cites Modeling context in referring expressions,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling context in referring expressions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.197496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.704689Z digest=sha256:99a14211b6c8a250143499e0764cc036cad17fd11edfabd33daad17232f0a1f4

Observation 1cfe4aa3-3e50-4d06-89ae-46bce9abedf8 · outbound

This paper cites Learning to count objects in natural images for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to count objects in natural images for visual question answering,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.182018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.709730Z digest=sha256:6aaa78599c0853740e66db094373c664e1bb098297b9c05a15c208174fdb3ded

Observation 5513a91a-6151-457e-b88c-d5b26a4e0e03 · outbound

This paper cites Deep modular co- attention networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Deep modular co- attention networks for visual question answering,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.166305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.714365Z digest=sha256:686437a6cf08b394f30eb3e7f0eac3092021deac9543025f7134be5a7ae8f308

Observation f5eae359-dab9-4bab-b5de-cd6fe415a587 · outbound

This paper cites Learning to reason: End-to-end module networks for visual question answering,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Learning to reason: End-to-end module networks for visual question answering,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.148630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.719073Z digest=sha256:67a331fb0123b0b0dddcd355e74f714f9a85408845b5e0f616b73d252ba2a1a8

Observation a20f6b27-33c3-4b36-a99a-bdcd5ada6b5e · outbound

This paper cites A simple neural network module for relational reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions A simple neural network module for relational reasoning,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.723698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.723698Z digest=sha256:009bdeec43511d190fe85ed60c4eb476907c3b9f3a2eccd44740bd53423967df

Observation 263fe6b5-0431-4cd3-9e01-869ad04fabc7 · outbound

This paper cites Inferring and executing programs for visual reasoning,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Inferring and executing programs for visual reasoning,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.120140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.728584Z digest=sha256:a1d337998223eda2abd2a81f71bd7165cfd11dc59d4e5ea4e4d5c192859762c1

Observation 3982b222-2158-42f5-9eaa-2b9aa7fdce41 · outbound

This paper cites Film: Visual reasoning with a general conditioning layer,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Film: Visual reasoning with a general conditioning layer,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.102590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.733087Z digest=sha256:b11683c25445003bdd6b40b650328075e79bc40b9f2cb054f4b0a9389db71524

Observation fe3807c5-9063-4fe5-b5ee-88a6eb62baa5 · outbound

This paper cites Ssd: Single shot multibox detector,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Ssd: Single shot multibox detector,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.086625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.737651Z digest=sha256:02f6b5fe3657fd6f277afdcbdb54172cca9284f326524e95960725178952a915

Observation cb4a4a63-7375-47e7-9e60-4868dba22dee · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-14T13:55:26.742259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:55:26.742259Z digest=sha256:e592573abd09c71af8f0451d8556d0cffb7afb3d1e3ebd08d146da0eee04d92a

Observation 113d83f9-c6f5-44b9-9627-4a54f7bf7bbf · outbound

This paper cites Referring expression generation and comprehension via attributes,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Referring expression generation and comprehension via attributes,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.069979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.747240Z digest=sha256:0fe5c5430c20193ae16d2c55b02ff758092c6fced7887bd9f8c3f37b88b042c1

Observation 056540f1-7b59-45d8-bbe2-4ec13864cf9e · outbound

This paper cites Modeling relationships in referential expressions with compositional modular networks,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Modeling relationships in referential expressions with compositional modular networks,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.052588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.752432Z digest=sha256:504ba6f1846ba940239c738d3ea728245948a67f1ccac8de4f17ade3a595b445

Observation ce81e276-226a-4c4f-9822-e6f343ce6cf9 · outbound

This paper cites Grounding referring expressions in images by variational context,.

Multimodal Unified Attention Networks for Vision-and-Language Interactions Grounding referring expressions in images by variational context,

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T13:55:27.036101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-14T13:55:26.757084Z digest=sha256:3edbbdc7ba89094430524d1631cdb9218422a2611e47d7b14912fa7129bcb087

Pith citing papers

Observation e801b09e-8839-475b-bff5-3d70b5134379 · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T12:22:25.662003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:22:25.662003Z digest=sha256:fa524faee19240b12dcc7112ef46d2d4fa7404cdc606786dfd029fb7d4a8b1f9

Observation 727d9b13-5265-4449-ad2c-613515bb85a3 · inbound

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement cites this paper.

ViASNet: A Video Ad Saliency Network for Predicting Dynamic Saliency and Viewer Engagement Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:15.881026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T08:50:03.872247Z digest=sha256:2427a11513ad755628b654da61a5a768210b9ddcc52b3575c51d5228bfa4ec6a

Observation 419bf000-2dcf-42f0-a8e1-6427010faba4 · inbound

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET cites this paper.

Alzheimer's Disease Diagnosis using a Multimodal Approach with 3D MRI and PET Multimodal Unified Attention Networks for Vision-and-Language Interactions

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:39:30.843132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T17:44:50.842397Z digest=sha256:3add86538b18ece3461025430d6bfe2f16918aa5ca03d523584fdb0fe59a95e6