Pith. sign in

Paper Citation Record · LEDGER

Microsoft COCO Captions: Data Collection and Evaluation Server

As of 6 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 100 inbound Pith citation observations for arXiv:1504.00325.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1504.00325 v2

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-12T21:38:19.467233Z

measured 146 of 146 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 100 of 145 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:46:59.069582Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact11
  • verified fuzzy35
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1634
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation d0fca574-faff-4a9f-8e20-7b730f51d70e · outbound

This paper cites Learning the semantics of words and pictures.

Microsoft COCO Captions: Data Collection and Evaluation Server Learning the semantics of words and pictures

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.689981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:81ffcf7c3cdc2f1957fbc61eaa3fc63f8698672c8a90b5a64b992c1a6a0fd44d

Observation 1c09e5c8-962e-4b4a-bc29-c5da351e87a7 · outbound

This paper cites Matching words and pictures.

Microsoft COCO Captions: Data Collection and Evaluation Server Matching words and pictures

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.698876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:e0bae5ae37690a0d593faab11bf6c6f0e4df5f94fe9fe6c9cd2960782abb8364

Observation c7b7e6de-7abc-4a3e-a224-150ba7f3c525 · outbound

This paper cites A model for learning the semantics of pictures.

Microsoft COCO Captions: Data Collection and Evaluation Server A model for learning the semantics of pictures

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.703779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:62cb96bf4de07474ecd917849cdbe9503c36b7f5bf383fd6761c0ccf9f1b9d76

Observation 57b70b4b-c22c-4850-8692-9e48499cb3de · outbound

This paper cites Baby talk: Understanding and generating simple image descriptions.

Microsoft COCO Captions: Data Collection and Evaluation Server Baby talk: Understanding and generating simple image descriptions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.708165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:82949250420f09b0cd03e0b3d7f022ed4114684809c78239760232f4bfab7fdb

Observation 82f6e1ee-94a8-4e44-b5f1-5c375839de19 · outbound

This paper cites Midge: Generating image descriptions from computer vision detections.

Microsoft COCO Captions: Data Collection and Evaluation Server Midge: Generating image descriptions from computer vision detections

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.712570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:6f4051e95818a5abc3fe8a2e54e8541e343a04e287dbc921599a760345dad847

Observation cb1a9009-dfe2-4684-8e31-80338c82d45b · outbound

This paper cites Every picture tells a story: Generating sentences from images.

Microsoft COCO Captions: Data Collection and Evaluation Server Every picture tells a story: Generating sentences from images

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.716746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:36ca00af7527917d36e5c3bc96bf9c47d369ec7b12a74bf7a2c1b3cb1c3e1b73

Observation 9f48f1cb-22cc-4356-821c-26e071949ace · outbound

This paper cites Framing image de- scription as a ranking task: Data, models and evaluation metrics.

Microsoft COCO Captions: Data Collection and Evaluation Server Framing image de- scription as a ranking task: Data, models and evaluation metrics

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.721122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:4595d33e90d357c0c3d7b622ae85438b342ab1d461b96e6ae7a498ccde60d297

Observation dc7b9a30-0ff2-4b85-a102-cd58173a3c7d · outbound

This paper cites Collective generation of natural image descriptions.

Microsoft COCO Captions: Data Collection and Evaluation Server Collective generation of natural image descriptions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.724924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:f0a73ac5a6d90adee2766ffecd85bb0fb64ad91f4c525ffd5f7ef2da2fd554d7

Observation 5c1c9ca7-f529-49d4-ae98-2f5e9703f1ce · outbound

This paper cites Corpus- guided sentence generation of natural images.

Microsoft COCO Captions: Data Collection and Evaluation Server Corpus- guided sentence generation of natural images

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.730308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:85910387a52a7c3614d55ea420832375e49689d7085e5c38074578ea90c5c681

Observation 41de0ff8-83cc-4e98-82c2-b7fca65b13b1 · outbound

This paper cites Choosing linguistics over vision to describe images.

Microsoft COCO Captions: Data Collection and Evaluation Server Choosing linguistics over vision to describe images

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.737094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:5ca206eb40784118a9a975e6c6b217749443a4abce7a5855752d1d63d5a9a6ed

Observation a4a31652-9f26-4fe9-b24a-e6bc5f82d9f7 · outbound

This paper cites Distributional semantics in technicolor.

Microsoft COCO Captions: Data Collection and Evaluation Server Distributional semantics in technicolor

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.741339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:a49adb79f329cbc52f848364f721aa0135b5677bbc020843d374ede8ce8c4565

Observation 8f655521-d30d-4f68-9c63-7510ffe7c672 · outbound

This paper cites Automatic caption generation for news images.

Microsoft COCO Captions: Data Collection and Evaluation Server Automatic caption generation for news images

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.592504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:2cd8116616c384fbc873685d8622fd6db0bc2d812c20301aad41958fd090a2a7

Observation ac2e3079-6ed4-4153-b2ee-838764b63ce9 · outbound

This paper cites Image description using visual depen- dency representations.

Microsoft COCO Captions: Data Collection and Evaluation Server Image description using visual depen- dency representations

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.597889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:36d91e2ee295018061e79f3eaa8431db333f38e746e310185561fbed22f1cd5d

Observation 98a7367c-97ca-4a7e-9948-cf5b97de07cc · outbound

This paper cites Deep fragment embeddings for bidirectional image sentence mapping.

Microsoft COCO Captions: Data Collection and Evaluation Server Deep fragment embeddings for bidirectional image sentence mapping

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.602473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:70022e8bea5eb50e0013f535f3aacc26a0ebb6604a01cd1a834d3cbba0752348

Observation 86889648-4499-48f6-9195-bc952ca2bd19 · outbound

This paper cites Improving image-sentence embeddings using large weakly an- notated photo collections.

Microsoft COCO Captions: Data Collection and Evaluation Server Improving image-sentence embeddings using large weakly an- notated photo collections

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.606466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:1ed628eb9fd2f09f607aad82a7c7d7fa231ccf37d2ec5d0474063402c50c3f1e

Observation 1ae5dd68-bdeb-43ef-b75f-6c524968aff4 · outbound

This paper cites Nonparametric method for data- driven image captioning.

Microsoft COCO Captions: Data Collection and Evaluation Server Nonparametric method for data- driven image captioning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.610330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:cb2fbe78965cf8eba6c1ad9640ac736423921789c637c9655a55a970b9b336c7

Observation f4542b6b-0a82-458b-a20c-774b2daa6870 · outbound

This paper cites Treetalk: Com- position and compression of trees for image descriptions.

Microsoft COCO Captions: Data Collection and Evaluation Server Treetalk: Com- position and compression of trees for image descriptions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.615522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:de6f6709c790d510159539b2cfb6a528345f0345b57d2b268944431234c221e0

Observation 291577fd-92a6-457e-9050-c3c088e71f00 · outbound

This paper cites Autocaption: Automatic caption generation for personal photos.

Microsoft COCO Captions: Data Collection and Evaluation Server Autocaption: Automatic caption generation for personal photos

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.619855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:40e424b1d2030a1a6fec704860990f7ffce655eda1a9c62e3e7f39d023829ebe

Observation 3cc59b09-c17e-4f3a-a93d-cf5eb152198a · outbound

This paper cites Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world.

Microsoft COCO Captions: Data Collection and Evaluation Server Is this a wampimuk? cross-modal mapping between distributional semantics and the visual world

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.624692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:1565ff6efcc77687b2a9437fea09ff739f4827db2feb11eaa983740794bde8d9

Observation 94234b57-8d52-44eb-bb38-a2c6ddb18bba · outbound

This paper cites Multimodal neural language models.

Microsoft COCO Captions: Data Collection and Evaluation Server Multimodal neural language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.628927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:8bf79bb559a8a6f857388e7a3981c3ec3d401b9aef0ca3693537f3aa8ab674c2

Observation 5c2f80e5-1a2f-4fcf-b129-baddddfc405a · outbound

This paper cites Explain Images with Multimodal Recurrent Neural Networks.

Microsoft COCO Captions: Data Collection and Evaluation Server Explain Images with Multimodal Recurrent Neural Networks

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.544943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:0e6a8f622410c874b8f41352dc2951b36aa528fb72cc3ad38ca9f40402fcc5dc

Observation 03732dc8-78fc-4552-9f8f-4866fea7b655 · outbound

This paper cites Show and Tell: A Neural Image Caption Generator.

Microsoft COCO Captions: Data Collection and Evaluation Server Show and Tell: A Neural Image Caption Generator

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.562091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:716b82aacf54bf60598723cf22a440d4824d8f3fdba921dc9fab5c1be21e88e5

Observation e7e304ae-2e03-43fe-9340-6ee2e07c86b5 · outbound

This paper cites Deep Visual-Semantic Alignments for Generating Image Descriptions.

Microsoft COCO Captions: Data Collection and Evaluation Server Deep Visual-Semantic Alignments for Generating Image Descriptions

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.570591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:01705a83f7bc2082532c93a8f4599a5d506e7c4eab277c1a282ec5f8fae858c3

Observation b9d6221c-6c39-4453-9ee3-03755662c2d7 · outbound

This paper cites Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models.

Microsoft COCO Captions: Data Collection and Evaluation Server Unifying Visual-Semantic Embeddings with Multimodal Neural Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.578820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:b00ff889f13312b7501e15d69b4655853302dadc1874a1852d35ad15e3efce8c

Observation 38a27fdf-cc51-4e93-aeca-4126a352feab · outbound

This paper cites Long-term Recurrent Convolutional Networks for Visual Recognition and Description.

Microsoft COCO Captions: Data Collection and Evaluation Server Long-term Recurrent Convolutional Networks for Visual Recognition and Description

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.586510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:4bed66831a37bdf7975dc38d47c2915f382956bd717c8ed4b2e53c00960c3eed

Observation efc3f455-2459-47b9-b689-70a1e5269383 · outbound

This paper cites From Captions to Visual Concepts and Back.

Microsoft COCO Captions: Data Collection and Evaluation Server From Captions to Visual Concepts and Back

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:50:48.615902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:775b8008fb4d6dd331d4003c02346d4890ee107e3aeb158e717e1a5c90b55e36

Observation d86d2490-9e03-4a9e-a5d7-3a9874a3396e · outbound

This paper cites Learning a Recurrent Visual Representation for Image Caption Generation.

Microsoft COCO Captions: Data Collection and Evaluation Server Learning a Recurrent Visual Representation for Image Caption Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:13:44.685711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:78c70ce2c03fc7f8b530e3a5834d24259bf0cdc277f91ae79b36660b723e33c0

Observation c907f701-2f57-43e2-aea5-024fe507cafd · outbound

This paper cites Phrase-based Image Captioning.

Microsoft COCO Captions: Data Collection and Evaluation Server Phrase-based Image Captioning

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:27:54.738193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:35306db747bacb32f0b029e8322ee3ed321a8efd823fced7ce4f0127efffade1

Observation de81183b-7de0-4072-bac2-9168a8ed40dd · outbound

This paper cites Simple Image Description Generator via a Linear Phrase-Based Approach.

Microsoft COCO Captions: Data Collection and Evaluation Server Simple Image Description Generator via a Linear Phrase-Based Approach

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:28:08.033637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:053fc385154762000fa5f5e820f14f3a0265c43f8f4d3aca349a806409ca392c

Observation 1d77adc9-3c10-4059-899d-bced179f77a5 · outbound

This paper cites Combining Language and Vision with a Multimodal Skip-gram Model.

Microsoft COCO Captions: Data Collection and Evaluation Server Combining Language and Vision with a Multimodal Skip-gram Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:22:41.881735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:3ef65d1f7d837263bab72dfa2b3a6abed01e752ac9a46e0e465bfe0a70137ed2

Observation 2f0568ad-78e0-42f0-9f7f-c3896dd6c5d6 · outbound

This paper cites ImageNet classifica- tion with deep convolutional neural networks.

Microsoft COCO Captions: Data Collection and Evaluation Server ImageNet classifica- tion with deep convolutional neural networks

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.632663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:6ba53196bfb822ff2443b57f160ce5c067d8870471dfced12950dccc8d93b625

Observation 428d2919-29c5-437d-a0f9-285d189e2cd5 · outbound

This paper cites Long short-term memory.

Microsoft COCO Captions: Data Collection and Evaluation Server Long short-term memory

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.636676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:5cdf8d1457f453cc499d21618a726d56324ec4b11a2e733ccb05e49d3a692ab7

Observation 33aff716-cbf6-49d1-a04c-c274e441e07a · outbound

This paper cites Im- ageNet: A Large-Scale Hierarchical Image Database.

Microsoft COCO Captions: Data Collection and Evaluation Server Im- ageNet: A Large-Scale Hierarchical Image Database

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.641273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:7bc43e1ac609ccbaf8c09e0d43a3c314e524b4feb3a711d240243b42f64bdfa9

Observation 4b252d52-dfa3-42a1-92b5-78623237dee2 · outbound

This paper cites The iapr tc- 12 benchmark: A new evaluation resource for visual information systems.

Microsoft COCO Captions: Data Collection and Evaluation Server The iapr tc- 12 benchmark: A new evaluation resource for visual information systems

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.645070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:6fa85d8e5f3e66df1653d2634d0867a811c9020bf08575723fcd05801bb8ddfa

Observation 4750ed89-1a4c-4fbe-b071-67072afa8a9f · outbound

This paper cites Im2text: Describing images using 1 million captioned photographs.

Microsoft COCO Captions: Data Collection and Evaluation Server Im2text: Describing images using 1 million captioned photographs

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.649158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:6547fee7278b56cf3c32d1eb0ea724f5f858d48006c0e686f4e3dbb7bc54db49

Observation 7653d2c6-a9dd-4254-9bac-275b183e821e · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

Microsoft COCO Captions: Data Collection and Evaluation Server From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.652571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:184bd5b7b1f054a545dc210e155ad08f428e9734cbcadfaa1e79d9634ffa872f

Observation 277e00be-035a-44ed-96fe-62a9c5efa986 · outbound

This paper cites D ´ej´a image- captions: A corpus of expressive image descriptions in repetition.

Microsoft COCO Captions: Data Collection and Evaluation Server D ´ej´a image- captions: A corpus of expressive image descriptions in repetition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.656751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:f8b5aad8646bb35d706dd185295761e7d1fd88030387a6116f922376ef652782

Observation f2aada12-4605-4731-86b1-874c5a757062 · outbound

This paper cites Microsoft COCO: Common objects in context.

Microsoft COCO Captions: Data Collection and Evaluation Server Microsoft COCO: Common objects in context

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.661376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:9b25f135c605d82f89f56b76b348e2d260242258f031e48c06a6d53e83c2b8f4

Observation 5d661a79-cba0-487d-83f6-da0887741bd7 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Microsoft COCO Captions: Data Collection and Evaluation Server Bleu: a method for automatic evaluation of machine translation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.665127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:042cac099fe13daa22e5b2e2ab3e613e180242dd50ef57ba7661d8ce414f9a73

Observation 9fad988b-6649-4ac1-8c8c-290c1edb26f9 · outbound

This paper cites Rouge: A package for automatic evaluation of sum- maries.

Microsoft COCO Captions: Data Collection and Evaluation Server Rouge: A package for automatic evaluation of sum- maries

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.669743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:415181f964e1c5f0b33994a892fd6613dca7bab6220e74b3365e1c9def02743c

Observation 7ecf7082-9ca8-453d-9202-242a72b1091b · outbound

This paper cites Meteor universal: Language spe- cific translation evaluation for any target language.

Microsoft COCO Captions: Data Collection and Evaluation Server Meteor universal: Language spe- cific translation evaluation for any target language

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.674306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:cb75dca2e3c30bb0183cd1a73954cebdc6bb95e954328e4f98a0f87a4f91b3cb

Observation 23314617-5ffd-4bbe-bf18-80548dae3b04 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

Microsoft COCO Captions: Data Collection and Evaluation Server CIDEr: Consensus-based Image Description Evaluation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.553333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:1454fa98e3543c22ae47514a2654e4978054a9882c7cf3895ac199cd2936af1e

Observation 085283e9-723e-4202-8f9a-9e8580326e13 · outbound

This paper cites The Stanford CoreNLP natural language processing toolkit.

Microsoft COCO Captions: Data Collection and Evaluation Server The Stanford CoreNLP natural language processing toolkit

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.677996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:59c59552b1e0e96362cd9ae6f6de84494f7b629b6b88e911c61b0ca51797168b

Observation da1ed1de-4fd7-4090-8531-4c881a5bad70 · outbound

This paper cites Wordnet: a lexical database for english.

Microsoft COCO Captions: Data Collection and Evaluation Server Wordnet: a lexical database for english

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.681870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:81d1b3fd8e35a2d6308cf6b769ba5630ead511de3daf1cdd1f7240c381c1344c

Observation db021f40-3aea-4ed4-9790-b371a21ff1bc · outbound

This paper cites Comparing automatic evaluation mea- sures for image description.

Microsoft COCO Captions: Data Collection and Evaluation Server Comparing automatic evaluation mea- sures for image description

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.685652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:2771da4cfa2a00a50819960f51b205272867882795fded43c8f15ff5a6271006

Observation 2176ac8b-8262-4b6d-a8c0-908ffcd2eb8f · outbound

This paper cites Re-evaluation the role of bleu in machine translation research.

Microsoft COCO Captions: Data Collection and Evaluation Server Re-evaluation the role of bleu in machine translation research

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-12T21:38:19.694427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T21:38:19.467233Z digest=sha256:1bd50d76139d82d63ed4d3035efd727f3e01aa888423c171335ac6daadfda061

Pith citing papers

Observation 1b224bc0-b1cb-4f54-b4c0-d64a0ee6fbf4 · inbound

Crowdsourcing a Dataset of Audio Captions cites this paper.

Crowdsourcing a Dataset of Audio Captions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-24T18:09:47.742982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T18:09:08.661697Z digest=sha256:22d68c24ee28dd985f7573983bb90721837890718dfba8146ef6fa269742931a

Observation 8beb0f93-e65f-4839-94cc-14a66091233d · inbound

VisualBERT: A Simple and Performant Baseline for Vision and Language cites this paper.

VisualBERT: A Simple and Performant Baseline for Vision and Language Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 100

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:59:37.780504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T23:59:37.636404Z digest=sha256:be1b350be12a064c72a8083cbb9916aa80f3e5396385a9ebb69c3ccee6c01790

Observation 28f8284f-880d-4a7b-af0e-ed6aecf21ed8 · inbound

#PraCegoVer: A Large Dataset for Image Captioning in Portuguese cites this paper.

#PraCegoVer: A Large Dataset for Image Captioning in Portuguese Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-24T13:44:32.039341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T13:43:24.358023Z digest=sha256:c1d050fdd9d4f8dd6bed54d50298bf5cadb286cb4600af2015010de6b4272539

Observation 1be85add-ca19-4722-b024-b84682139dba · inbound

Image Captioning via Compact Bidirectional Architecture cites this paper.

Image Captioning via Compact Bidirectional Architecture Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-24T12:19:27.034831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-24T12:18:47.846727Z digest=sha256:048e6c0331bb9206ed49a84521a22f8a343f2b72e8d29de8cee88eff9f3f0d80

Observation a5e5b6e1-e1be-45c7-b0f9-9aec249a0b21 · inbound

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language cites this paper.

Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:50:00.679782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:50:00.546571Z digest=sha256:4cfd02ea7ed95c915eb2e980db8cef0bd4c9eed53734a18b43064dcbbac812f9

Observation 9bd9d5a2-0df3-4e36-a224-5533aea59dae · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:d1fca90845f8575176648cc5f8fa033648a28b62aacc7ee786cca1f813c632f0

Observation 85ff7a0c-afc8-45c8-9c53-afc8e6b0d871 · inbound

CoCa: Contrastive Captioners are Image-Text Foundation Models cites this paper.

CoCa: Contrastive Captioners are Image-Text Foundation Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-15T10:53:08.487450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T10:53:08.292063Z digest=sha256:af9593405c5ba9513d278ee8815f9c1dcdab52d2aa8367386d88fed17683ff70

Observation 5063c6e6-0e76-45d7-91e6-dcb9abafb7a7 · inbound

A Generalist Agent cites this paper.

A Generalist Agent Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:24:49.989850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T06:24:49.833638Z digest=sha256:39645a416dd42749edd60663caf35dbe719d5f46ab5a70a3faba68d6a52b072c

Observation 4972a877-902d-4b91-a06e-e0cc5db38a2b · inbound

PaLI: A Jointly-Scaled Multilingual Language-Image Model cites this paper.

PaLI: A Jointly-Scaled Multilingual Language-Image Model Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 146

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:29:06.050752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-16T09:29:05.956863Z digest=sha256:d7be44739c74aace6be471950f86d0d18386df5e567450d4bf22c06bb5e8ba7d

Observation dcaf4239-7be8-442b-998d-32601fd7840e · inbound

Learning to Detect and Segment for Open Vocabulary Object Detection cites this paper.

Learning to Detect and Segment for Open Vocabulary Object Detection Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-24T10:26:08.085742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T10:24:52.653023Z digest=sha256:f489c82b5324bcb7b606ba63eec0d0a66de48a67c1417fecaf71dc9d30579f65

Observation 72eb375c-a89a-417d-b9bf-3a5a0d6ec89a · inbound

Sigmoid Loss for Language Image Pre-Training cites this paper.

Sigmoid Loss for Language Image Pre-Training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T13:05:36.533326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T13:05:36.460932Z digest=sha256:7cefbc32a445dd49cf99cc1ec0cb785e096a094465b6e8c738f406063cd3648c

Observation 8ec80962-ecf4-4b0d-ab78-4a128e7c5401 · inbound

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality cites this paper.

mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-24T09:04:15.032403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T09:02:31.211260Z digest=sha256:99f664be92301a69f9a6dcf61b65f3f12a89ffddcc05d19415ca7bfcda5476ae

Observation 6b5bb468-a0cb-45db-aaa3-856fe30870c9 · inbound

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model cites this paper.

LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-15T08:41:04.775361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T08:41:04.743886Z digest=sha256:753f8b8af7061a14032d76d48c366eddcff0087661f10a33b4f5d2faa818352f

Observation d65ca8d9-1a4f-4151-ac49-c1be7d63de2f · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-15T02:43:47.935048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:33689a87a93bc49dd95fac4e8402c740299aaf2972e2066175498086b6579090

Observation e3d6e35d-5704-4e6a-80c4-1fd9563f1deb · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-13T23:30:00.510458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:76e8b421b0c54a51b8a7a64a625094508f6f606e3d41652aee8e1c9c26063a65

Observation c5faac31-01d7-4ec8-a997-7452b3a0043f · inbound

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration cites this paper.

AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-24T08:29:11.333273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T08:27:35.798991Z digest=sha256:7ce8c60617ad98177b470fef05b9d512a5d10c74771afceaab08499accf97bbe

Observation d72dc9d0-68d7-48fa-9da9-00833f1d5585 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:cb8a6372fd98961b7bc55d7cd780f531d6274eaf54079e54e4d967c6c2302216

Observation e0f64b52-267f-4c9c-b3b2-2ea512a05ab7 · inbound

MMBench: Is Your Multi-modal Model an All-around Player? cites this paper.

MMBench: Is Your Multi-modal Model an All-around Player? Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T17:20:53.687692Z digest=sha256:a5f97e5b0a78c711d8b01aeb4b22c8583beabe896768b3ee816e7a7e67796848

Observation e9e2fc6e-7849-44a6-92a1-1935d9dbbb30 · inbound

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models cites this paper.

OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-14T01:52:01.402804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T01:52:01.163900Z digest=sha256:f871a965302212e3d48bbf79b6278fb3b4281e9b2334560b02f8e0ed1531d52b

Observation 58ace55f-b7c3-46a4-9f8f-c0bbfadc5816 · inbound

Directly Fine-Tuning Diffusion Models on Differentiable Rewards cites this paper.

Directly Fine-Tuning Diffusion Models on Differentiable Rewards Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T09:11:32.066723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T09:11:32.018262Z digest=sha256:b1b8e48508ae6e30f6ff921c755c0564f33321ead28cfd1f45d19829584a5e3a

Observation b73ff6d1-a6a1-4f16-8cef-e628fa8411e4 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-15T23:26:06.540494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:94fe978d93a5fe70fd4c1c4610bda81c17565578359d4a53031adef3f391d503

Observation 2c1c9cb5-18d6-4fad-a99b-b5568e944a1b · inbound

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions cites this paper.

ShareGPT4V: Improving Large Multi-Modal Models with Better Captions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T17:08:12.795765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T17:08:12.727773Z digest=sha256:d9ff7efd6d5290ad5d7c4356ab9dd3a7e842cbdd046afdd6430b1358cc828055

Observation 78707856-7a45-4c03-941d-9e768b6f2e47 · inbound

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks cites this paper.

InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-13T22:46:09.829246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:46:09.693156Z digest=sha256:261ec027bcc33baa57f78eed435c4fc960db6589e9f71c4774210d76e56a7e50

Observation 2a143420-8309-4e39-ad92-46cc99f2d939 · inbound

DoRA: Weight-Decomposed Low-Rank Adaptation cites this paper.

DoRA: Weight-Decomposed Low-Rank Adaptation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:26:21.591591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-15T22:26:21.449135Z digest=sha256:e20017dde64a3ed22b36cb9dcce4eeae7f7e7b107e6b1c6868e08d85eeae5310

Observation 634ff84b-25a1-48e8-8fe9-f4856ffe6db7 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T04:09:36.145523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:86d86946f3c6b5186fdafe815a579cec80e389f6fbc7fb61ad15682371390040

Observation 1c5c2acd-fa89-4835-89a3-b54106d45d9d · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-17T07:44:47.588449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:e80a90957b28ceb8c148e50baf8d69a56c7f1d6a4ccbc12db90cd133310f3188

Observation 6eaba2df-25ca-4931-9e93-1bbf2b57fd07 · inbound

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis cites this paper.

ANCHOR: LLM-driven Subject Conditioning for Text-to-Image Synthesis Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:43:42.944905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:42:10.041654Z digest=sha256:4f6db51cfa82d8a06ee581549f24b5ad7f8d8e474b457321a015e94dedfabdf5

Observation 7ae65d8e-6ae0-4fe9-b52c-9b6763fe6d40 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:cb4433e6403939c95dcd01ccc808a6df0c7123100560a0c638e2ba8b8d237257

Observation cde83335-67ae-484f-a674-e202b2d35f41 · inbound

A Survey on Vision-Language-Action Models for Embodied AI cites this paper.

A Survey on Vision-Language-Action Models for Embodied AI Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 232

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:25:54.539820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:25:10.150459Z digest=sha256:911b03fecf150ef98760447ae0f965b00c9a83faffede89aa70c9bd1e1450c24

Observation 56fcaa07-b58f-47c1-92a5-af4c13df9346 · inbound

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference cites this paper.

PipeFusion: Patch-level Pipeline Parallelism for Diffusion Transformers Inference Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-24T01:18:42.384966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-24T01:17:11.261301Z digest=sha256:51c208d913a693a59be5bee038875bca5c74097929f6c5c14b103409a02ebcc6

Observation a907fd74-2589-47b4-b26a-14e1b92d53ae · inbound

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models cites this paper.

Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 133

Resolution
verified exact
local_arxiv, observed 2026-05-18T06:38:36.868269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-18T06:38:36.517935Z digest=sha256:be3a715b90acb1fc97944f7e8fb19c4b425c36f684fb2bc0b064162297b42863

Observation e27abcfe-0a82-4689-8963-6ac137d7c451 · inbound

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models cites this paper.

Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-15T01:55:12.582208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T01:55:12.501409Z digest=sha256:58712f26b6ae40b582a263e298939e2c3d45f7a68452cf74ea550cf00ec0ebd0

Observation 6a2041fe-417b-4c19-9a2d-0547cea3490b · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:97400262cce0382131c81bcc3fd1c3c3962d644e9a6b28737a76eec35d32a6b7

Observation 3613bcb3-16b4-4e07-9744-43832d64649e · inbound

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation cites this paper.

Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-15T22:09:16.420682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-15T22:09:16.001309Z digest=sha256:abb1d3e6e2a540e68b6e6066d3aa24edcd6a5cce152ee708427a830ccbfec7db

Observation 849f954c-b5d3-45ed-9ad8-a1e4fb5f5c82 · inbound

Adversarial Hubness in Multi-Modal Retrieval cites this paper.

Adversarial Hubness in Multi-Modal Retrieval Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-23T06:42:39.767194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T06:39:36.039613Z digest=sha256:919e5d2e8ffb707b6d4a010ac6a4c38322dd710608df2ba7dcc23c1e6a220e32

Observation b421cf42-446e-47dc-a096-f07fe94d7417 · inbound

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation cites this paper.

Speak-to-Structure: Evaluating LLMs in Open-domain Natural Language-Driven Molecule Generation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-25T08:15:33.999096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-25T08:12:15.133695Z digest=sha256:dd8c6e5fe05dfa27787bec77aff477d5fdc45bd20eda890710deccc51bcf778d

Observation c6fb3883-0635-402b-9e8c-7b33de0352f1 · inbound

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP cites this paper.

Grad-ECLIP: Gradient-based Visual and Textual Explanations for CLIP Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-23T02:22:25.136680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-23T02:21:09.704462Z digest=sha256:3c0e2c9d3f3ac10279adee89df1cbc7b57b9afaec92f6a0e785b9bf42446d296

Observation 056870e0-0de0-41d7-ba02-e8b606b04876 · inbound

A Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions cites this paper.

A Woman with a Knife or A Knife with a Woman? Measuring Directional Bias Amplification in Image Captions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-22T23:55:15.074377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T23:54:15.135595Z digest=sha256:9042be76f6920cb8f6fd00cfc18d904326feeddaf2865b9eb1c878287f045db0

Observation ef1c1f6c-8326-4732-9178-5854e652d087 · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:22:12.291375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:716a3b80eabc5735202d8c6acfe89a4f7987d4a28953f142c74c32acfa1988e5

Observation 75e1d32f-7ffe-4658-b2fc-ef8ea69bf78d · inbound

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization cites this paper.

$\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.977492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T18:02:23.305313Z digest=sha256:d41a9ffdf07abbfb79df8eb300ac59588e39cd20b22419fb82b6893f9a6515ec

Observation d9765233-7f4f-4ef4-a691-6772fcad41cd · inbound

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback cites this paper.

We'll Fix it in Post: Improving Text-to-Video Generation with Neuro-Symbolic Feedback Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-22T18:56:58.326343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T18:56:39.735334Z digest=sha256:1195c535d43a427416597c7be9e8ab815e89cf04357c827bd7100bcd3454f7c8

Observation b62dd826-55a9-4bed-bf3b-221e700e048a · inbound

Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition cites this paper.

Masked Language Prompting for Generative Data Augmentation in Few-shot Fashion Style Recognition Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-22T18:05:00.839252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T18:02:59.412368Z digest=sha256:1f2ba5c48233740e0f477f935c04471876f7b5d426cfae964d3476d05ece4283

Observation 32f12def-dc00-4938-9830-98e138ab2a0f · inbound

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models cites this paper.

HyperCap: Hyperspectral Land Cover Captioning Dataset for Vision Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-22T15:16:43.738611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T15:15:19.523035Z digest=sha256:34bee07ea54f090afcc604bd8afda072fdf796d167a923e29b5ef5781b950801

Observation 071ea3a2-2e7b-4f78-babf-6cbf29e18ff0 · inbound

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM cites this paper.

Slot-MLLM: Object-Centric Visual Tokenization for Multimodal LLM Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-22T02:10:56.220111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T02:06:35.204166Z digest=sha256:630f35467b1e55f85585a49f4fe0aa62e589382fa35e0ef4e411d75f0e858e41

Observation a69c3971-c78c-4fb8-8cb5-9ff221c9e252 · inbound

Common Inpainted Objects In-N-Out of Context cites this paper.

Common Inpainted Objects In-N-Out of Context Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:42:15.974706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:38:44.069681Z digest=sha256:f99e1802c88780c7a098fee52e9d8c940c57076bb2aca1f3aa0b769906685674

Observation 70a55e50-6593-4259-b46b-b15de2234537 · inbound

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments cites this paper.

VS-Bench: Evaluating VLMs for Strategic Abilities in Multi-Agent Environments Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T11:57:16.282786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T11:57:08.314088Z digest=sha256:ea92edcc79a34646f804d5648137508350629c37d2bbc7350c0b0a15d3f45122

Observation 2ed96b91-ddab-4159-9191-dda38d3ba10b · inbound

T2UE: Generating Unlearnable Examples from Text Descriptions cites this paper.

T2UE: Generating Unlearnable Examples from Text Descriptions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:59.069582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:59.069582Z digest=sha256:c34db8785fe6abb059a7ec27a242a93452e8847efbd422b0e731cf73ce53c784

Observation 104d7e7f-843c-4d54-b4ea-67aa9a54e84a · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 116

Resolution
verified exact
local_arxiv, observed 2026-05-19T00:36:56.333655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T00:34:38.247099Z digest=sha256:d0169aef090f7d16eb36cc6b78a89daa77be285a4a4fb9ad3f1ebccad9396d20

Observation 6a2dce26-9059-4f6b-8fca-e1f536e47088 · inbound

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models cites this paper.

Decoding the Multimodal Maze: A Systematic Review on the Adoption of Explainability in Multimodal Attention-based Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 116

Resolution
unresolved
no resolver link, observed 2026-08-06T00:02:58.992894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:02:58.992894Z digest=sha256:d3d42bfcaf91ed63be6f3e560073225c6bf418d594eff852ba685847975497fa

Observation aaf5043a-1208-4a05-9244-460064c6cb40 · inbound

Adapting Vision-Language Models Without Labels: A Comprehensive Survey cites this paper.

Adapting Vision-Language Models Without Labels: A Comprehensive Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T23:17:06.035332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:17:06.035332Z digest=sha256:d36add2201cbc619c6c6223992d95833ac9308a027485b9705a15ed2942600c0

Observation da60b3e5-d993-46ec-b9e0-2c35daee4b04 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.657104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.657104Z digest=sha256:4d8bcc045407a695d56727751181bd3a5dae365aa037b5c788333cdcc7249c54

Observation eb61d106-f26f-4c75-bb70-9a412e1885f1 · inbound

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models cites this paper.

Segmenting and Understanding: Region-aware Semantic Attention for Fine-grained Image Quality Assessment with Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T21:52:47.542522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T21:52:47.542522Z digest=sha256:c1f196a894d88d78eb27e5faff3d7722b1a7ab9ba547364606fbfddb6eb874bf

Observation 04e72c13-00da-4d94-94e1-a2a106e8c979 · inbound

BERT-VQA: Visual Question Answering on Plots cites this paper.

BERT-VQA: Visual Question Answering on Plots Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T20:36:40.251887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:36:40.251887Z digest=sha256:4a4215a09476f2858b53e641fc1d826b8fee01d6a79a5adcaaf7465d00b18020

Observation 4e04d4c3-2a9c-4b4a-96f7-cf27e24fb2fc · inbound

MobileCLIP2: Improving Multi-Modal Reinforced Training cites this paper.

MobileCLIP2: Improving Multi-Modal Reinforced Training Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-05T14:59:19.506021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:59:19.506021Z digest=sha256:572041b9abbcab976dde912f25ffb3c43cc0f7835cffd5097842ce98aef0659a

Observation f9d83025-6180-41d2-896d-ad21f1ed39d8 · inbound

Aesthetic Image Captioning with Saliency Enhanced MLLMs cites this paper.

Aesthetic Image Captioning with Saliency Enhanced MLLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T10:15:50.303995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:15:50.303995Z digest=sha256:3399194faf3bf7d4eb15eaf4d7d96f24ea52e401bb954523435db811b7b1b0fb

Observation c62b17c4-81a9-43fd-884c-47acded4c33f · inbound

RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness cites this paper.

RT-VLM: Re-Thinking Vision Language Model with 4-Clues for Real-World Object Recognition Robustness Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T12:57:49.943388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T12:57:49.943388Z digest=sha256:8c67878fe01791fcd2500bd1a1f09aeb708107bde2b218481ef5c86677251423

Observation c4d6bcf3-0373-476b-a9ec-42fd2926a9e3 · inbound

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs cites this paper.

Towards Meta-Cognitive Knowledge Editing for Multimodal LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T05:12:44.083105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:12:44.083105Z digest=sha256:8f2130d55b6ba942a45ed9e6448166559dc98d1a316ed26d520760de1f6447e5

Observation e0573306-2c83-4580-99cc-0589aa918643 · inbound

Multilingual Vision-Language Models, A Survey cites this paper.

Multilingual Vision-Language Models, A Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-18T13:02:37.444440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T13:02:08.000814Z digest=sha256:f3bf67e0ceebd84d652e2876a51b1548e2ec0db0e2e08ba3c7db575a81150f68

Observation cfdadec1-1bae-4b93-9db0-91c600144b44 · inbound

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models cites this paper.

Evidence Recomposition and Predictive Context Residualization for Visual Attribution in Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:58:37.449797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:58:37.449797Z digest=sha256:6fd22e0da2c67a4f7f6756353c34a06415ccf36fcfbfc7c08d816d8d9583294a

Observation ce35e764-dc80-440c-9624-75d245551088 · inbound

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images cites this paper.

Locate-Then-Examine: Grounded Region Reasoning Improves Detection of AI-Generated Images Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-18T10:16:14.416463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T10:13:08.112072Z digest=sha256:fcc58359430bd32479e98cfc1800d5590000d0329dc9d380d3197a1b202c81d5

Observation 43c6531b-7c27-4a43-86d7-ad1e0e6896d5 · inbound

Activation Steering with a Feedback Controller cites this paper.

Activation Steering with a Feedback Controller Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T21:50:41.390967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T21:50:05.186376Z digest=sha256:3e461d5f8b64dd11c10430f6d62b215fa04b774c179b1797fe41e911b9600462

Observation c1f49450-7e1b-424f-b3e3-caf108aba6aa · inbound

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects cites this paper.

CaptionFormer: Unified Segmentation, Tracking, and Captioning for Spatio-Temporal Objects Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T09:31:55.130506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:31:55.130506Z digest=sha256:7967365773b504eebda593b5c2cf53817da2f74180559b2059c09895acf64672

Observation 4b131054-0fb6-4782-952a-2473e4857ac5 · inbound

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback cites this paper.

StableSketcher: Enhancing Diffusion Model for Pixel-based Sketch Generation via Visual Question Answering Feedback Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-18T05:30:55.340950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-18T05:26:42.034972Z digest=sha256:bd37ee6fa16be8fb3622f23f9605024643d78cb22384ec6e702c74728a025813

Observation 3a55d6f6-ec61-41a2-a2b1-845abd6728ba · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T19:54:20.234562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:d558179c9ab23a04845b354229d4fb4375a9e17cb455ebe1bed35d067f22c748

Observation 6a66a247-ca63-4b4d-9fb9-2a60e134f0f1 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T21:43:58.707091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:43:58.707091Z digest=sha256:3d71c6c64bccbcb08f18511202b3175df350c08d3b6d6ef82eea1186c551a473

Observation 5c40efce-1f1f-45ce-8a00-1988d3ee5d9d · inbound

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation cites this paper.

Semantic Router: On the Feasibility of Hijacking MLLMs via a Single Adversarial Perturbation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-03T20:27:03.812410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:27:03.812410Z digest=sha256:be099f6ab80483093ad8d7aa16f14845e01321e38901a02557bf7524fcb2b627

Observation 3bc9f3e4-c916-459b-be60-ac19a6b184fb · inbound

Benchmarking and Enhancing VLM for Compressed Image Understanding cites this paper.

Benchmarking and Enhancing VLM for Compressed Image Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-25T07:26:41.749325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-25T07:26:06.192538Z digest=sha256:ef49970961dd971054e34b4c3b6d3fa924f3f1bdfea7d0fd319e9dfdc6092b63

Observation a19db48b-6f61-4ad4-b66d-8bd77e2769b2 · inbound

A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization cites this paper.

A Geometry-Aware Efficient Algorithm for Compositional Entropic Risk Minimization Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T05:17:24.815840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T05:17:24.815840Z digest=sha256:57ee84ac59a5014d3e8832bd19a66f5986b58be2bc1e9afcbf14a08e308a7adb

Observation 861bfd8e-0120-4558-bb20-82a55f20ac88 · inbound

Xray-Visual Models: Scaling Vision models on Industry Scale Data cites this paper.

Xray-Visual Models: Scaling Vision models on Industry Scale Data Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T22:25:58.434888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:25:58.434888Z digest=sha256:e6b9ed417f25ca33f130ede9416234d4d03c658a66c52e0e3cb09b11d0e7ab7c

Observation 1654ccdf-30af-4f0b-9e10-79da33a60132 · inbound

Concept Heterogeneity-aware Representation Steering cites this paper.

Concept Heterogeneity-aware Representation Steering Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:46:04.720421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:46:04.720421Z digest=sha256:6fc0fe66bf6ae2cf26bc8159fd4d4df3867ef66d16f0883c7168a19cc14563bd

Observation b3eeed85-7c5e-4003-bfbd-110ea924877e · inbound

LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation cites this paper.

LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T22:58:23.958105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-13T22:55:49.617588Z digest=sha256:8a192966176930c9eca625b37cd304d98ea7497d83d4906e01eb980b4dfcbc53

Observation bf121e0a-cf53-49e2-a98d-dc675e30b284 · inbound

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward cites this paper.

Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment Reward Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:59:19.379119Z digest=sha256:1248f17a2bdb114fa2a9abe7d11f53d8b1e504fa1edbc334da3c556df0eb420c

Observation c68ea044-6281-4f65-b0f2-d487fc3b967a · inbound

Batch Loss Score for Dynamic Data Pruning cites this paper.

Batch Loss Score for Dynamic Data Pruning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T19:17:42.229302Z digest=sha256:d6ff5626d8c5dcc4418a2488736d34e14a8db8d3dccdbf29fde739371a183337

Observation e86717f9-2227-49dc-9280-22bdebb26ccf · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:9576f2e8f4379be24ba9f82be788eac599971ed447ea55655e4ebb5104e76397

Observation 0f8eb8c4-a5a0-4ee5-9983-29578d8f499a · inbound

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment cites this paper.

Vision-Language Foundation Models for Comprehensive Automated Pavement Condition Assessment Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T17:30:39.410040Z digest=sha256:c17e1e9af38843f9d81c516b9e8d148fb17dd2ea0b4eee8fa941aa0612d31619

Observation 682baa85-5976-4d35-bbf2-5c2796df9f60 · inbound

Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation cites this paper.

Tessera: Unlocking Heterogeneous GPUs through Kernel-Granularity Disaggregation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:50:39.600498Z digest=sha256:a7d6d8d34a0c3e55c8df8c84562d50a8f1596b428f27508fafa24eb8dcfdaab9

Observation c7cbac4f-55ca-4d26-a927-767a880de252 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:737e015d7a01a5be3e3ccd8241fb8ed8babc874bdbc80cba44593694df9d826c

Observation 6eb3bc23-8ff8-4672-b700-2f559dca0f0d · inbound

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment cites this paper.

TIPSv2: Advancing Vision-Language Pretraining with Enhanced Patch-Text Alignment Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T16:00:12.466258Z digest=sha256:bceeecb894901f20d4f8a5e6d2e09b3e08b41d78b0144f5cfd0e2a32698b1c2e

Observation c1a64b82-9a6e-45ea-89ea-66d3ba2eafec · inbound

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks cites this paper.

Challenging Vision-Language Models with Physically Deployable Multimodal Semantic Lighting Attacks Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T15:58:28.202606Z digest=sha256:4996592567eea91b8fdcc625b0c3f59c7352fd24b6a2e818df3f01f862a6b0f6

Observation 72392037-408c-4038-b241-f0825ca65924 · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T08:18:59.432486Z digest=sha256:f127e8dc8eeda8847d15dd0ebca6742ea0078e5e0a62c04765d86f3fee05692e

Observation a7e2dc5f-ca7d-49c0-8e81-b447cb4febae · inbound

S-GRPO: Unified Post-Training for Large Vision-Language Models cites this paper.

S-GRPO: Unified Post-Training for Large Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:53.670309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:53.670309Z digest=sha256:ad8e04d24ddd752c4a00b2f6e211fbe9c7c8f24b76b39bc0735ceb1ded1c56b8

Observation 9ca71ed9-5063-4ca7-a91f-bd613812408d · inbound

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning cites this paper.

GaLa: Hypergraph-Guided Visual Language Models for Procedural Planning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T06:26:06.208036Z digest=sha256:6e8137e68b7fd1d40ce05735638afd1216d4289ab1f8c66b553a1ccbe8718257

Observation 30411ee3-4c6b-40b3-acf4-2c89db9e798d · inbound

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models cites this paper.

From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-10T05:44:37.891638Z digest=sha256:7717872f69976e1fe3044286f049f9a2356deb1e4176800d3ecd03c2e7cc8e51

Observation 0d3419f0-c575-48a1-bb60-db89462d6476 · inbound

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models cites this paper.

VLA Foundry: A Unified Framework for Training Vision-Language-Action Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-10T02:10:04.003151Z digest=sha256:f3413b33eb8ff069a9e1901467c35ff7b666bce7926b5ef18e14a76a0d211830

Observation 9db6f033-ac28-4632-8ede-a143d374c9ae · inbound

Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection cites this paper.

Exploring Hierarchical Consistency and Unbiased Objectness for Open-Vocabulary Object Detection Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T08:41:12.351655Z digest=sha256:e8d72192cac0a75b5f6d7ab52bead3c26bd62ff52a32ac54089cc500c0d2c04f

Observation a707248a-0b80-46c7-98f0-6f61ee59764f · inbound

EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure cites this paper.

EASE: Federated Multimodal Unlearning via Entanglement-Aware Anchor Closure Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-09T18:29:15.893882Z digest=sha256:ef294c8701f0e605c1418d2ede5cb17471a721959058b275fb916cab643d909d

Observation 56655c96-ac24-4294-83f0-8cc0be789acf · inbound

Statistical Consistency and Generalization of Contrastive Representation Learning cites this paper.

Statistical Consistency and Generalization of Contrastive Representation Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-09T16:55:04.086365Z digest=sha256:152427ac4dec0f0194fe2bd48a64e2e7e5c5f7b5236fcdd5ee552882935a40bc

Observation 440091d4-9a67-4776-b93b-7481e51b04fa · inbound

Statistical Consistency and Generalization of Contrastive Representation Learning cites this paper.

Statistical Consistency and Generalization of Contrastive Representation Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T09:04:04.649210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T09:00:00.110353Z digest=sha256:2881ae4bc249111b8824a029d3ccccfba634706c35fe18b9063ea69b40c413ad

Observation ca5676f7-4456-4ec2-83bb-05227f70caed · inbound

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning cites this paper.

Sentinel2Cap: A Human-Annotated Benchmark Dataset for Multimodal Remote Sensing Image Captioning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T18:17:35.846352Z digest=sha256:9fbb38ab9907b16a7091c5b2c227a11ad733a5a0229220d6ff9d1d61d798820c

Observation 869a73a3-2b58-4301-909d-fbae7dd547e7 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:409a0139ef0b802bc0d446a030f10508db87ed0c72b55ddecf9c32a0f1e2f529

Observation 7341c44f-03db-4fb8-8821-41a6c14ff47d · inbound

ZAYA1-VL-8B Technical Report cites this paper.

ZAYA1-VL-8B Technical Report Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T21:38:19.742717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-12T01:15:16.607346Z digest=sha256:70f35762bdb5b33026d29a434a8c6b813f19d4f47e0cddd5920ed6cb4f36fc12

Observation bbfff318-c287-4577-8b86-c56bccd91a9f · inbound

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models cites this paper.

Learning to See What You Need: Gaze Attention for Multimodal Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 123

Resolution
metadata mismatch
local_arxiv, observed 2026-05-14T20:22:56.233042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-14T20:13:18.813131Z digest=sha256:67ef68c9c73b6623ba88acc78b77ea7eb578121b5035ba8f098c79cb32c3a1be

Observation 9bb89455-6644-4877-865f-7f36bb5ce8a6 · inbound

OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems cites this paper.

OxyEcomBench: Benchmarking Multimodal Foundation Models across E-Commerce Ecosystems Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-14T02:18:37.559140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-14T02:13:53.218695Z digest=sha256:b62dc459d0c2ad4b6e0318c128c021010327920fc8f7ebe2b5b1559b44b407d6

Observation a2483a73-46e5-4985-8421-ec4436a5fb34 · inbound

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations cites this paper.

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:23:37.723191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T18:19:05.137641Z digest=sha256:85ec60a5a90e3645eb9463579eb44021ca4ad86514ebdc71f81858cdda915bcb

Observation 4bda0e1f-e7df-4e18-bfa9-129e3909eda8 · inbound

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations cites this paper.

Right Predictions, Misleading Explanations: On the Vulnerability of Vision-Language Model Explanations Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T19:15:00.870582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-30T19:08:14.197339Z digest=sha256:bfec8225d529b7cd61a0b804a4aa80b038a70516d2097c539fa3c8c5d8b03708

Observation 4849b43a-6994-4b13-ba5f-672e3a3ce30d · inbound

DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers cites this paper.

DiRotQ: Rotation-Aware Quantization for 4-bit Diffusion Transformers Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:52:48.408408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T21:48:35.203469Z digest=sha256:1f4ad5ef392b4b0f508fac7c20f5154e37cebfa6a3e770467ff28f81afad3e3b

Observation cce6b67d-dcbf-4033-af31-f2a2a54d8585 · inbound

DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models cites this paper.

DarkLLM: Learning Language-Driven Adversarial Attacks with Large Language Models Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T18:33:37.936267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-20T18:31:48.770507Z digest=sha256:a25ccb1b28508485ffeec6116ba96f4193d29a1c398049108142b1a55e603900

Observation 877641bb-11f5-4ad2-8da4-1b4fd1c4bb75 · inbound

Modality-Decoupled Online Recursive Editing cites this paper.

Modality-Decoupled Online Recursive Editing Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:49:53.972582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-21T08:44:56.543284Z digest=sha256:f0a2aba4b2f54358fc57fb89d00a48fca80835702712c4318ab775d4cd38e9cc

Observation c96ca966-6bf5-4cb7-846a-c483f92a05ff · inbound

Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability cites this paper.

Measuring Cross-Modal Synergy: A Benchmark for VLM Explainability Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-22T06:06:08.766897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T06:05:34.039507Z digest=sha256:80edd3d408ab5236500c5e139d86fb32693dda268325c733ff4537f93f2b7165

Observation 6bfa0228-402d-44e1-bc8b-95fe8119fbb3 · inbound

Cambrian-P: Pose-Grounded Video Understanding cites this paper.

Cambrian-P: Pose-Grounded Video Understanding Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-22T05:51:08.961660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-22T05:47:13.493461Z digest=sha256:ad328dee417c7b111bde54c15bbdcf931722f63005128c0b5e5fb8a5137caa99