Pith. sign in

Paper Citation Record · LEDGER

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

As of 23 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 29 inbound Pith citation observations for arXiv:1908.02265.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.02265 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:54:41.910373Z

measured 74 of 74 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T05:43:42.788616Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1675
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b8fae9c0-6b9c-4390-b0be-a5f30ffae3f0 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:54:42.527937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.715981Z digest=sha256:4a1e1a712ae5b94479d2cb84a59ce7d4d5735fad1b2080569684ed6bb8d8e715

Observation 1fa016a7-422d-4cba-8bd5-0e8e5ff667e9 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.721129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.721129Z digest=sha256:b800726fe5181a81d09dd70bd3f57abd23c027ad004f04824e4bdc3236c48724

Observation 405e9517-e5a1-4f83-a976-2fb3e3f4af00 · outbound

This paper cites Lawrence Zitnick, and Devi Parikh.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Lawrence Zitnick, and Devi Parikh

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.505305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.725405Z digest=sha256:1c4fafb2e6d34ee6b411f513277e96a96ac44dac9e35bc0ef852b9f333431909

Observation 7a95221d-8e15-4709-b46a-7c8810a23519 · outbound

This paper cites an unresolved cited work.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:54:42.491585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.730077Z digest=sha256:7739f40a725c29b267c87f5bd31e52042af91048cc7c15fd1e99691f2ddbb4d5

Observation 3c67ad70-a369-49f9-99d6-58153b0cd641 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.734742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.734742Z digest=sha256:fe7c8b99dd721602cef39e70e0c2ea0a23aebf47e385d57a7c635a57307ec7fe

Observation e75e8e42-e1ab-4f93-aa03-248e64581f24 · outbound

This paper cites foil it! find one mismatch between image and language caption.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks foil it! find one mismatch between image and language caption

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.477685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.739613Z digest=sha256:e1d8e60f4c5556a949b82dd7ff229c8032c1d23f1c7054218cc556381dde51a2

Observation 27ab7fd0-c77c-4ec6-93ce-0865b1f768ca · outbound

This paper cites Embodied Question Answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Embodied Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.744340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.744340Z digest=sha256:8252aba218b2d2c9da7be4670655581ae27a3c167d2130d2d1f10eca8ab2acfb

Observation 3589cdaf-866d-4745-9e10-7d0d480c7bb9 · outbound

This paper cites Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Vision-and-language navigation: Interpreting visually-grounded navigation instructions in real environments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.455218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.748498Z digest=sha256:83b5883aebb25c062602fc2ab004880334e49ec50a30dc80a000b2e1be726c1b

Observation d6e2e915-46ed-455d-947e-f9c7b94ce0ff · outbound

This paper cites Don’t just assume; look and answer: Overcoming priors for visual question answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Don’t just assume; look and answer: Overcoming priors for visual question answering

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.440601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.752931Z digest=sha256:9133b17b161befb200dae9cf3c586781edcee2a16827976e53ab0bad0edae549

Observation d98e4f41-c2bc-4011-8377-39a3080dbbd4 · outbound

This paper cites nocaps: novel object captioning at scale.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks nocaps: novel object captioning at scale

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:54:42.048363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.757219Z digest=sha256:b1ff0498ac440f391f8be16c3249973e070ae0a3423b487350a2d03b7ea0abd1

Observation 6c76b031-dcbe-45a3-a7f7-899775b61797 · outbound

This paper cites Deep residual learning for image recognition.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Deep residual learning for image recognition

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.761689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.761689Z digest=sha256:da81ea35db2f54a40be4cf8750790a0efeb7a04ac91faafbcbcf86f0e811f7e2

Observation 729e08fb-c372-4219-8db9-5453e054af65 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.765854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.765854Z digest=sha256:b6f1f71f225d5494be0af97ee8ba8258678603617030d37cc452fc3a066488fb

Observation bede3cb3-c9f0-4199-8aea-471431a0a59e · outbound

This paper cites Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Peters, Mark Neumann, Mohit Iyyer, Matt Gardner, Christopher Clark, Kenton Lee, and Luke Zettlemoyer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.417597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.770843Z digest=sha256:227bfae2e18e151692857ac41ce4e4f37c73522ceeef492298340431d3d0d6b4

Observation c4977e9b-8f1f-483f-8838-3a155826a247 · outbound

This paper cites Improving language understanding with unsupervised learning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Improving language understanding with unsupervised learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.403348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.775087Z digest=sha256:6dfe3095bdb1e339ee469e44171fb77be5e9617be0a006bd6fe5ad4be984eec2

Observation fd8d8339-eb10-4cec-a6a3-60aa5660ff73 · outbound

This paper cites Berg, and Li Fei-Fei.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Berg, and Li Fei-Fei

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.779434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.779434Z digest=sha256:77042c03301548a5f7b60f7766222c345fdccb0aa04f8821dd43724af05c6e3f

Observation 44c8df1b-c7ff-47d1-ae94-4e94ac6e09fd · outbound

This paper cites Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Visual Genome: Connecting Language and Vision Using Crowdsourced Dense Image Annotations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.783733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.783733Z digest=sha256:d8c3dfbadf186c746e7e03eed5785639bfdfc5671fe98a08fd9bc6e5b8044891

Observation 4a931e69-6555-4cb6-bef2-a8d7bb5b2719 · outbound

This paper cites Aligning books and movies: Towards story-like visual explanations by watching movies and reading books.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Aligning books and movies: Towards story-like visual explanations by watching movies and reading books

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.379199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.788156Z digest=sha256:5e90e669aa0e7bdf5ddabe1cb1ed63104acf99cdf849cb563457b8a1d908786b

Observation cc52f0f3-3bb0-4761-8ec2-e57fad484d82 · outbound

This paper cites URL https://en.wikipedia.org/.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks URL https://en.wikipedia.org/

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.364622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.792364Z digest=sha256:fff040732c8fd00fd48c14a419fe5365f734dca701f13ce194e4502befea8ee0

Observation 3d237df6-a9d9-4524-b513-dd37c38d4592 · outbound

This paper cites One billion word benchmark for measuring progress in statistical language modeling.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks One billion word benchmark for measuring progress in statistical language modeling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.349631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.796867Z digest=sha256:42d359053313343f2c7192a2700d4b1722801e6738ec0e91d5264ff5507bc7ab

Observation ea437fef-cd26-4a3b-8feb-dfc3f86daa18 · outbound

This paper cites Colorization as a proxy task for visual understanding.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorization as a proxy task for visual understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.335581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.801205Z digest=sha256:cb9f94cba0dd0cd2ca745aba0185a1978bb1263130d20faeeca06aff829d928e

Observation 18dcce10-bc79-4977-9c3d-3252c8e6af42 · outbound

This paper cites Shapecodes: self-supervised feature learning by lifting views to viewgrids.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shapecodes: self-supervised feature learning by lifting views to viewgrids

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.321415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.806094Z digest=sha256:90cf54c561be7cfbaa572ad1231108eb3199c9865ca42bb5eeac2394caac7854

Observation dd8d5424-451d-49a5-b858-18ead5c46d8a · outbound

This paper cites Look, listen and learn.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Look, listen and learn

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.307014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.810443Z digest=sha256:59a23fc659d67bd57ab28a51408b8d1eaa577592597298bffcea09fc64a1c052

Observation 7f047bb8-390d-4482-904e-be35fe57c304 · outbound

This paper cites Learning features by watching objects move.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning features by watching objects move

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.290416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.814544Z digest=sha256:0d070d913699fd8c7707a62f483b72493942d34e4fd6a5468c25a00a1ad8b2ae

Observation 99e38133-ad65-4215-9531-35d2080f0e34 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.818952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.818952Z digest=sha256:6a5c97170a61d2b39bc256e38ff2e3b8fe3c04d9d45059b8848f4f4ca330f545

Observation 5cebbe19-81d2-4467-a158-ebb002609938 · outbound

This paper cites From recognition to cognition: Visual commonsense reasoning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From recognition to cognition: Visual commonsense reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.822905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.822905Z digest=sha256:692db50e8c41020cda8673d7f199fd797fb9d031fcd7b26b5f791ff7384bba1f

Observation cd8edebe-f4bb-49fe-a9e0-129127aa1104 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.827134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.827134Z digest=sha256:1d167e011cd414576170117faf9b109d7a69c5f86bac5f7b48b2ac82a0642f74

Observation aecd10d1-5847-46e2-9aa3-ba19f2aa93cd · outbound

This paper cites Attention is all you need.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Attention is all you need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.831210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.831210Z digest=sha256:3306c59eddcd87832f744263b64d7c76dc2fb53e1f236c711a565d484184e650

Observation ac8f1a2f-9a10-47d5-bc4a-d0afdfe82cf4 · outbound

This paper cites Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.835507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.835507Z digest=sha256:6f8600ddfc7940aa524ade861725630cc1f17cec56a9102c731e1a80ef5fc2bf

Observation 0e96ad61-628f-4e10-8e03-2e6cf31489c0 · outbound

This paper cites VideoBERT: A Joint Model for Video and Language Representation Learning.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks VideoBERT: A Joint Model for Video and Language Representation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.840774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.840774Z digest=sha256:633e4ccb8982b8ea05acf04b4861a7d7b6921c319ac788e2d9928f355c8323cb

Observation 14aab524-892e-41c8-affd-f3078120b476 · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Bottom-up and top-down attention for image captioning and visual question answering

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.845410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.845410Z digest=sha256:27e1c82319dde7439655819bc695f52fdbe077e7c4621f1a719ade5e8e1ea526

Observation 1fb83a2d-8745-4d42-9ed8-88462481760e · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.230555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.849617Z digest=sha256:ccf013dd70ad22bcfd259cf46d92de8a87dc15d1dee1a2f99d17fa65bbd60c2b

Observation e857b7d4-7934-4897-8e8c-1273306c134b · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Referitgame: Referring to objects in photographs of natural scenes

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.216875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.853703Z digest=sha256:92361de34ff158eb09da5580d8dad0b96642231afa6935dbd24ddcc7fda0e07b

Observation 7dad33ac-6be7-455a-9c84-8837ddbc74ae · outbound

This paper cites Mattnet: Modular attention network for referring expression comprehension.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mattnet: Modular attention network for referring expression comprehension

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.202206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.858106Z digest=sha256:a480b8c53a65cde46cdf2c18263a925f57fe9626041fe97f8beadc76863f9a76

Observation 6a66df2f-075a-4e3c-b7ba-c46c540614d0 · outbound

This paper cites Mask r-cnn.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Mask r-cnn

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.186388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.862141Z digest=sha256:b0dfd2666646f7eb7713184e323a54f28efd4323d7b935066c54fe373f8efd5c

Observation 7e59628b-d5b8-49b3-8cee-a0b6dfe14169 · outbound

This paper cites Stacked cross attention for image-text matching.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Stacked cross attention for image-text matching

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.172389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.866210Z digest=sha256:a85bf10a84b4004ed1ddbf745622aabb6416ecd83b03e7fd8b0f54e0f297ef97

Observation c77166b2-c5d1-4e60-9b1d-7334467c6ed2 · outbound

This paper cites Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Dynamic Fusion with Intra- and Inter- Modality Attention Flow for Visual Question Answering

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.870680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.870680Z digest=sha256:e1c53e035d4933ea1420bdc575fbba39dfd6169bd6ddc5711b86137c07dd4a06

Observation 411ea18a-47e1-432f-b2df-a10478bbd2db · outbound

This paper cites BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks BERT has a Mouth, and It Must Speak: BERT as a Markov Random Field Language Model

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.875322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.875322Z digest=sha256:889d5df3953340680216a24733d4596d1c2d0d747c416ce00d99787f76b8eed1

Observation a1486f96-6943-4e04-b33e-8eddb4b02df5 · outbound

This paper cites Unsupervised visual representation learning by context prediction.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Unsupervised visual representation learning by context prediction

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.156813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.879689Z digest=sha256:db77bf8e9ea6fc73639a44c51d91e2a12679144b76dc90eb9a13d3215a0c9e8d

Observation 71a9d6fa-5b50-4038-8fb8-8bb3039410a6 · outbound

This paper cites Colorful image colorization.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Colorful image colorization

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.142864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.884569Z digest=sha256:5bc38810064ffeeca5fdd16d91056173b5c238c81584d16b652c2ceddae3cbd4

Observation e46e166b-584a-4b27-81d6-e61180a2d4b8 · outbound

This paper cites Discriminative unsupervised feature learning with exemplar convolutional neural networks.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Discriminative unsupervised feature learning with exemplar convolutional neural networks

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.128503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.888833Z digest=sha256:070e7ed71ea997194aa8ce638cda35c48db36c2813edd0391f9c7b37097b786b

Observation 6be4d9c9-cc7c-43c8-8403-8b064124c98c · outbound

This paper cites Context encoders: Feature learning by inpainting.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Context encoders: Feature learning by inpainting

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.892878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.892878Z digest=sha256:ceca4fca0ceef00214c9a89dcb638dd0179a505339d53e8ed1ba982619c558db

Observation 0e43164d-2b37-4ba6-9ef9-e83ad7cfeaa0 · outbound

This paper cites Learning image representations tied to ego-motion.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Learning image representations tied to ego-motion

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.105716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.897142Z digest=sha256:6fdb5bd841dd984bcd7571d498b92ec65f3258d8e14adfe4645388a0d13bb001

Observation 90800b48-2a7b-473a-ae9a-2bc9413ba1a5 · outbound

This paper cites Shuffle and learn: unsupervised learning using temporal order verification.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Shuffle and learn: unsupervised learning using temporal order verification

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.091551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.901553Z digest=sha256:a00de9bcf178c280935fa115252ec5c2ba99fdf7f028b81fc2f2caff5c8ca9cf

Observation 1418ecdb-7534-4fc7-af88-9bd735f8adba · outbound

This paper cites Cross-lingual Language Model Pretraining.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Cross-lingual Language Model Pretraining

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-14T14:54:41.905995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:54:41.905995Z digest=sha256:8b2b44fa19ef8d7e7a946cff37b27d7b771fa6715dc23459309485bc3c69e6fc

Observation bfcc0816-f708-4cf1-88f2-24e2ba26c6a1 · outbound

This paper cites Courville.

ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks Courville

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:54:42.077360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-14T14:54:41.910373Z digest=sha256:1edec6557fa5860ed1bf4d3c9ffffd553566690b007551b4fb2662f74f5b2056

Pith citing papers

Observation ff96a687-5a8e-4fbb-971e-7ba9a2fcfec8 · inbound

VisualBERT: A Simple and Performant Baseline for Vision and Language cites this paper.

VisualBERT: A Simple and Performant Baseline for Vision and Language ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:59:37.786807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T23:59:37.636404Z digest=sha256:1c5a10f5e42e7724661ac682edb6ebfece38d1a62bf2fd55c487a7906b1f24a4

Observation fe831b02-d01a-4554-b47e-7eee89b55662 · inbound

Multi-modality Latent Interaction Network for Visual Question Answering cites this paper.

Multi-modality Latent Interaction Network for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:10:16.637749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:10:16.637749Z digest=sha256:ce809bbe09a5fddc6263a6d21d34bb92bac98619ca0117d5b51ab3992fdad2cb

Observation 22ca41b3-de21-4c32-925f-6933a0f7f51a · inbound

Fusion of Detected Objects in Text for Visual Question Answering cites this paper.

Fusion of Detected Objects in Text for Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T13:30:37.367264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T13:30:37.367264Z digest=sha256:3e76976fc1e5a959e175ae5112db364438106bf52bfc31d365adafed98b8b4f8

Observation 0ee6b0e0-0cef-45be-b6ad-9436b121f399 · inbound

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training cites this paper.

Unicoder-VL: A Universal Encoder for Vision and Language by Cross-modal Pre-training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T13:01:13.266000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T13:01:13.266000Z digest=sha256:3bc3d826398da0f232893842ca8d8cc6828c18d5cbb527a0eb5b32932a8e1f60

Observation 18ccb43a-3cd4-4824-9e2b-eec14122377e · inbound

LXMERT: Learning Cross-Modality Encoder Representations from Transformers cites this paper.

LXMERT: Learning Cross-Modality Encoder Representations from Transformers ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:22:25.590947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T12:22:25.590947Z digest=sha256:6216c1fbd53389141087cbb24371a9567c99e33c21bc2f33f4b3e1ecb6647d0f

Observation ea78c704-dae8-4dad-b4a5-705cab9935d2 · inbound

VL-BERT: Pre-training of Generic Visual-Linguistic Representations cites this paper.

VL-BERT: Pre-training of Generic Visual-Linguistic Representations ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T11:42:19.107745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-14T11:42:19.107745Z digest=sha256:6b6f596c740fd825ac4c7819c539ce34e503d1aa48bebf076ebba71030a8f9c7

Observation dde256fc-9f6d-475a-9589-44e54120cb74 · inbound

Text and Code Embeddings by Contrastive Pre-Training cites this paper.

Text and Code Embeddings by Contrastive Pre-Training ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T19:24:12.008090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T19:24:11.907204Z digest=sha256:4c7e223cb201dd4c35f45483ba0353506b9a31fd9e83e0a06bc14f316d4512da

Observation 2afcd137-29e2-48ca-ab18-d7b7ebbb7240 · inbound

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms cites this paper.

A Comprehensive Survey on Visual Question Answering Datasets and Algorithms ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-12T18:54:50.223310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:54:50.223310Z digest=sha256:6a86dbc341c475233e1b6960b3222ef6c38f5ade5de83ee0293a1804c77d9cf3

Observation 33a53c64-c86b-403a-a2be-a5762bb1a7d3 · inbound

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models cites this paper.

VisGraphVar: A Benchmark Generator for Assessing Variability in Graph Analysis Using Large Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T14:52:57.089001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:52:57.089001Z digest=sha256:2ce4e549b754905ac971396b8b060cdcd3bd6b2d105e25260fe7f1c7e6194efc

Observation f9d4fabe-2f81-400e-8c38-39e7c695f6e8 · inbound

SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment cites this paper.

SentiXRL: An advanced large language Model Framework for Multilingual Fine-Grained Emotion Classification in Complex Text Environment ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T11:31:54.596220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T11:31:54.596220Z digest=sha256:f98083bf901436e08fbeb3813e04bd3b9316d9ad2c9989410088c03c92d4c984

Observation c76cf31f-1383-49d6-8c21-87bd74ed90a1 · inbound

Multimodal Multihop Source Retrieval for Web Question Answering cites this paper.

Multimodal Multihop Source Retrieval for Web Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-10T21:42:24.403171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:42:24.403171Z digest=sha256:6afa8d531d30f561bc5d05aa78f88bf443fdfcc4ce5cdd2b996095988033461a

Observation d91dae2a-3456-4a2f-ad03-b7eae883678c · inbound

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models cites this paper.

Language and Planning in Robotic Navigation: A Multilingual Evaluation of State-of-the-Art Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:47:58.978874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:47:58.978874Z digest=sha256:62d0550a383a102eed64aa1d89ae699bdb9cc90bb3d09b511203454690065b15

Observation 883a7e11-56e0-4ad0-bde5-c29ede8dd36a · inbound

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering cites this paper.

The Quest for Visual Understanding: A Journey Through the Evolution of Visual Question Answering ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T20:51:53.740176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:51:53.740176Z digest=sha256:03aebd77ea9bb17d116127a355fa7b98741ca2e75ee9d37af61e2629e34c4e73

Observation 5b875199-5513-4653-99ad-c25625dfa6a5 · inbound

Performance Analysis of Traditional VQA Models Under Limited Computational Resources cites this paper.

Performance Analysis of Traditional VQA Models Under Limited Computational Resources ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:10:38.323443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:10:38.323443Z digest=sha256:67497f2645904546e5da2a9e6648a1006252f6b2aad579731f0bccf283a01144

Observation 87ea1a9c-bf48-4b70-a47a-df80579b30b7 · inbound

Vision-Language Models for Edge Networks: A Comprehensive Survey cites this paper.

Vision-Language Models for Edge Networks: A Comprehensive Survey ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-08T12:20:08.315864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T12:20:08.315864Z digest=sha256:2d5a29a708299ff78509d165f55ede9ddcd1e08fdb29cb11712253d7a1f01028

Observation 2f6bb367-4920-4ce2-9860-cea82207a6a0 · inbound

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI cites this paper.

Enhancing Surgical Documentation through Multimodal Visual-Temporal Transformers and Generative AI ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T05:43:42.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:43:42.788616Z digest=sha256:4608cc675e0c87fc51c0e312d6e77661e0ec0ab3bb71c9c8912f5b1f65e4ba2f

Observation 16d433d1-594f-41b7-90f4-22b6da9fb43c · inbound

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models cites this paper.

Detecting and Mitigating Hateful Content in Multimodal Memes with Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T04:55:28.198790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T04:55:28.198790Z digest=sha256:317396fcb376f60b5cd5bff84b7794d1c8833d11890fc150778d93de87e15dd3

Observation 18d9656d-e6ea-45f0-9bac-d593c441cf48 · inbound

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval cites this paper.

Representation Discrepancy Bridging Method for Remote Sensing Image-Text Retrieval ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:59:40.942765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:59:40.942765Z digest=sha256:58c4826fab59baf34e1d419513329231b1c7e7b66b3228a88b1885c30cf49132

Observation 8519c36c-eef2-429e-aa49-9d5a12ae840e · inbound

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation cites this paper.

LaVi: Efficient Large Vision-Language Models via Internal Feature Modulation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:13.773647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:13.773647Z digest=sha256:8d66a8f39d824f18351fec2baff929757d7d4a871bdc302d25e440502b0c667f

Observation 18a56846-eb1f-4f9e-8a8b-53be205dc05b · inbound

On the Resilience of Underwater Semantic Wireless Communications cites this paper.

On the Resilience of Underwater Semantic Wireless Communications ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:48:13.121047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:48:13.121047Z digest=sha256:16ac15a2d9d9f06dbd6700f56381d16dcf5df4e29bd97b27f1c8c4a8c7bfec56

Observation 233de5ba-3820-4e13-b61c-6c4c11479083 · inbound

Representations in vision and language converge in a shared, multidimensional space of perceived similarities cites this paper.

Representations in vision and language converge in a shared, multidimensional space of perceived similarities ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 6241

Resolution
unresolved
no resolver link, observed 2026-08-06T12:21:22.540405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:21:22.540405Z digest=sha256:35342448fe903e4ba2469563cbcea3f9a43d69bc8696c797cba68557954ac89f

Observation 691bee43-3208-4384-888e-71ed7341d524 · inbound

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation cites this paper.

AME: Aligned Manifold Entropy for Robust Vision-Language Distillation ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:20.512268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:20.512268Z digest=sha256:55c8e85efeed57da2c708bd5ea1a48d6e0b8e767d7582a9684f43f8c0b18214e

Observation e585f74d-cd12-4a2c-a89d-fea58f28007f · inbound

From Image Captioning to Visual Storytelling cites this paper.

From Image Captioning to Visual Storytelling ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T10:27:39.329615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:27:39.329615Z digest=sha256:3990fc052089ff6c02edd86375e4d00b9dbb5a459300e513fc8fa8ae8ac1cdcf

Observation 96d59a51-e9ae-4059-99ad-9723eee7f085 · inbound

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images cites this paper.

Appear2Meaning: A Cross-Cultural Benchmark for Structured Cultural Metadata Inference from Images ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:05:55.502242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T17:49:25.227568Z digest=sha256:799ac8b2ef2d915095113766cf6da4b8493f96b9242b2dd8c1d4571d204925b4

Observation 09885d11-81f9-41f3-bb38-1cec11183d43 · inbound

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments cites this paper.

Machine Intelligence that Understands Visual and Linguistic Information and Interacts with Humans and Environments ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-06-30T17:44:57.769037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T17:40:33.082748Z digest=sha256:513186e524062de4e05f882e29b8cd8571d9c9b56d9e102e5719bd537c91fede

Observation 4dfd39e3-fc51-48b1-9eb4-04611f821c0f · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.805103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:34c20926d61b6ee3813f03e97d9f1430b2a6a2163f991679762cd6209e747daa

Observation 878188b0-7386-4b73-9b97-e531469b0240 · inbound

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction cites this paper.

KG-TRACE: A Neuro-Symbolic Framework for Mechanistic Grounding in Antimicrobial Resistance Prediction ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-26T01:58:54.073345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T01:56:35.760918Z digest=sha256:c5e3dbc74b5c38a69c55ea5e544cc26c297dd778e9f8165a638c40db07c3f22f

Observation e1deed52-e357-4079-9b7e-30134d9c6f08 · inbound

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset cites this paper.

A Comparison of Fusion Techniques for Multi-Modal Human Activity Recognition on the HARMES Dataset ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-29T04:33:06.216866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T04:32:39.678440Z digest=sha256:56f0daf8bce7a655d6596ce5fa21921581507dbb9bbc747e238f0b95f9eeb732

Observation c2c49c0a-a86a-4b2d-82cd-01515858c4fb · inbound

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery cites this paper.

When Does An Extra View Help? Adapting Single-View 3D Reconstruction with Extra Imagery ViLBERT: Pretraining Task-Agnostic Visiolinguistic Representations for Vision-and-Language Tasks

Reference 112

Resolution
unresolved
no resolver link, observed 2026-08-12T00:26:30.762901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T00:26:30.762901Z digest=sha256:c00c0942fae1fe4dcf314fa199c0ea455b3eea7226ed184ab7bb2eef324f807c