Pith. sign in

Paper Citation Record · LEDGER

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models?

As of 22 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2502.06600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06600 v2

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:58:48.251828Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 83f9ae22-6664-4e3e-9ae7-facd58bf7ce4 · outbound

This paper cites From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.107848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.107848Z digest=sha256:52a44018a915e05f95b6bbb9cdc37345d0ae9184fe8fd5de8ea931d93bca2b7b

Observation 8d4fc923-288d-469e-8913-2e82bf926c33 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.703743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.112701Z digest=sha256:cf157d671bf8df55ecec9f7cae3f808707055e80749c173a7e73a490932edda7

Observation 6b138ef1-d3f6-4927-8fbb-ec2a0c8e06e4 · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.116638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.116638Z digest=sha256:35a647c389bd04ef0bc5a7ca3ada631631bb9fa491365079f2c05f790394d104

Observation 4e69acfb-352d-46d3-80bc-8845f2411477 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.693345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.120954Z digest=sha256:4bf33c00d9ec04ef4675d86a15beed4ec53e9e42f4d04a9e01321eb48dc96d84

Observation ab46c83d-b001-4dd1-8c64-eeadb081bf10 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.683627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.124736Z digest=sha256:48336436004458a63ddf6bbb385f14da3fa20ada47ab30fbaa7c3be023a64393

Observation 58f9782a-809f-4fa1-933c-d12dff79cff4 · outbound

This paper cites Evaluating Image Caption via Cycle-consistent Text-to-Image Generation.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Evaluating Image Caption via Cycle-consistent Text-to-Image Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:58:48.351746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.128379Z digest=sha256:59d52581772468d079feafbf6ec6017cf19d4b13e50b5e26933e69c5cf8cc757

Observation 1c9b6edf-f166-431e-9e0e-68708a488229 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.673976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.132436Z digest=sha256:0d0b5a017252987c839db16b5e8b128e699b727879e97aa35decf0386a5ae3d9

Observation 7d6d3166-20b1-43c9-ad06-3b9f14f4478f · outbound

This paper cites mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? mBLIP: Efficient Bootstrapping of Multilingual Vision-LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.135855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.135855Z digest=sha256:0448d6a698c7f70f39af11b3c9c09448167e64d6d979b6e92d3ee187e88341ee

Observation 054bdc68-448b-4a6b-900d-c853408c8bd1 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.664055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.139766Z digest=sha256:c2baf961e35631a34e2173a52574a1a683a8634c9b17d53222e6b3981d8380cd

Observation dcd3ac4a-2f04-42d5-86ba-ed55c2335154 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.653900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.143162Z digest=sha256:efacec46e2b96dd841c701a5381919d8cbe08887447f26d0837772bae69a68fa

Observation e8b7e6cb-d43e-4f60-bd21-d1e3534c6313 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.643350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.146463Z digest=sha256:3bc5be7263c2f396c911edb0e4d81d7cf5597dda4aa5c978412d16ac09e1b436

Observation f4cd076b-50e7-4c45-88e4-127b8d59cb26 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.632497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.149925Z digest=sha256:719e1ad40e9efe18b809d93928279aa0d840dda4a148d85844a79dcc73e4e17e

Observation f86b738c-e88b-462f-84b7-616c8fbf42da · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.622305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.152564Z digest=sha256:38c2350aa21d61ddef8ae0e38a61766fd69a3fca08ccc5a50300c8749ee158d6

Observation 0509c59d-0ce0-4f90-b056-583aad488a4d · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.611322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.155242Z digest=sha256:258915bf21e00f0488328a09b89e719b7ab56ea6943ff6c87919f77cfa4f1147

Observation 262e7b40-e123-4ea2-a546-18028f092554 · outbound

This paper cites FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.157795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.157795Z digest=sha256:609b0af75001c9ee4cd3c837f36f2eff9cc6a747941a12019901c2b365a8b905

Observation e94ee4ea-f377-4b0f-8732-a7ba9c3bb5ab · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.600724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.160659Z digest=sha256:7b4b1ec2acf7ab1d4a9857a0e5b8abb816f632b42494be666cd96f5dfeb20586

Observation 5dd98dfd-ec45-4b48-b16e-190a8b5ee264 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.589632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.163384Z digest=sha256:1f7fc0e081be69a30023147ab37d46ae2ac7aeef1cb8a9dc677f90bda06400f5

Observation 2a499ef0-c798-4326-958d-05772ebe347c · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.578526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.166210Z digest=sha256:531b9b916c0566e9cdefc1aec0e1a90c223ca39300b712592d865275f536583b

Observation dcabe155-8dab-4e76-91c0-868983d29dd0 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.567322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.168871Z digest=sha256:8ac99637410eb34f17424ba3494b3dd9ff72c27302113871e0025115b7041d3b

Observation 52e30fe1-49a0-4c7a-a590-48eef4db58df · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.556817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.171937Z digest=sha256:2671d71c6f1bcfc408ecb66813a1f31abd53ae6d89c237300211d9a25adeb693

Observation e827993b-975c-49bd-b84a-0efffdec7d0e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.545931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.175223Z digest=sha256:9e65d2aeb94f9e5f101b5d886bc42884764fc6ab096cd282becf28df63ed82a5

Observation cd7d2f62-94a8-4d85-a4a6-1f32bfd8f481 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.178520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.178520Z digest=sha256:957d71d91094a41d90964e9021181530e3fe1ea73a4cf368a6d0f1c440bdaaea

Observation 720e87c9-ba5b-4ca1-9252-aee9ddfa6290 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.529185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.181934Z digest=sha256:2061899db23f5cbf81c2b717ce55fc614a54395aa568e3b153fbbc6720ff1311

Observation 2c560504-c6a2-4af7-a220-b923ba8bf3ad · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.520419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.185359Z digest=sha256:164980a177618d2ef131634ae189c060bfd91f1cfceb85294479a5c318bcc83f

Observation 91ae5dad-3086-4095-902b-c279a8d85fa4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.511567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.188730Z digest=sha256:1973579d60c3886790a7f0e96032b4c3fb22ba5c77770bd492be16d03d84f39e

Observation 93864a7b-81d6-4f6b-9c5d-fe781abf2d33 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.502102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.192388Z digest=sha256:01cbb5e57d01f7e3a5cbdbdee89892b7aeca87807b082343ad207c5dc4db279d

Observation 3ff3180b-db37-4e66-a805-e6ba78042af0 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.492448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.195667Z digest=sha256:db56b23aa7bcae867a6c85439ec76c0cfc7bde411b40255241f61ebc4426fe6b

Observation 0128b44b-877f-49ea-b537-4d2c28710619 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.482354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.199242Z digest=sha256:0c914d3031684e872bc72c801fc150be9369e6eac9b4ae1e022bf0f7cadb9fda

Observation f4db2265-c362-4c07-bb56-8c486bd1e2b1 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.472530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.202653Z digest=sha256:d8ff51c3ada645ab5cd871022ce7a91c27f132061d76e1ee1e981b9ffce11c1a

Observation 88661f33-1a84-4df6-8210-b8a6f49236f4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.462397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.206031Z digest=sha256:b168da0de140165ae0035f2e6cb56b5a89a87747759872294ce2e10ff914539a

Observation 70bab504-fbe1-4dcd-94b0-de0d35278fbd · outbound

This paper cites Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Positive-Augmented Contrastive Learning for Vision-and-Language Evaluation and Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.209337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.209337Z digest=sha256:73eb42495e95bb22d0c6134edc10b0c25e9ff667b114552d60de0e654b378ece

Observation 30ead303-1f8d-4752-8115-bc8cd0becdc4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.452291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.213093Z digest=sha256:534a277c74e8b2cf43cc2162241fb2a982247d8c80d65e4d1eb8e63870e1486c

Observation a4d74aca-a523-4c46-b549-d18939fdc182 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.442298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.216626Z digest=sha256:e8fb9772103eb184aed4c3d6b66785f162327d5a43eb21e721d5e874d72e3f7d

Observation 91a1aeb2-7791-4bb0-806e-49a13e7c416f · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.431991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.219896Z digest=sha256:95635e9833157fda29117adfbf57e177167c35123fa0f8dc556fa70cbec5399d

Observation 5444d6e4-1a51-4718-8f98-e699647f718d · outbound

This paper cites G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? G-VEval: A Versatile Metric for Evaluating Image and Video Captions Using GPT-4o

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:58:48.309144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.223313Z digest=sha256:911a4b0f547b07429e83563cff27fc3aac2521fb6a9bdb3350e78d175d05418e

Observation 80b7e105-15a2-437c-83b2-7746e24bfdb2 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.421535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.227281Z digest=sha256:d407ff6d1a536ed7a57fc8c4f6dc384b95ffffd909f633077dfe0f81fdd71da3

Observation 8e4b798c-3d69-4815-acbf-c5c5db183cf4 · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.411384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.230725Z digest=sha256:a5296419b66b51788c510cc6c389031bb395a9f3769c702ebe8aff620fb59bc0

Observation 9fede902-8fa7-4281-a23a-e75100d5d95e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.401688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.234045Z digest=sha256:b47c7cc287f6945ae99bc7ec1c8f8d60beb092e7969147c8e500d85ec0cfb125

Observation e0a843d9-4994-4c1b-8a96-0b8dd8c8be83 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.237478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.237478Z digest=sha256:9068dba2723ff435c5ede9d7f69b89dbf70e9b6c65ce9a58e631bb8358b77879

Observation 3e996417-493e-4955-a628-82ecb0dba27f · outbound

This paper cites When are Lemons Purple? The Concept Association Bias of Vision-Language Models.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? When are Lemons Purple? The Concept Association Bias of Vision-Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.240959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.240959Z digest=sha256:e5555f261a0a42c57504513fef66e4ad713fc52b49cc28d3254fa89beefd509d

Observation 55d6fdf1-22a0-49e7-b462-f7684214929e · outbound

This paper cites an unresolved cited work.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:58:48.392417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-08T14:58:48.244606Z digest=sha256:e17e286b5b3c973c33f0fcb653884e204898c9b7420760b32d856403ffaf3577

Observation d455749d-bc74-4cf7-b9b3-4288580f2fab · outbound

This paper cites online" 'onlinestring :=.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? online" 'onlinestring :=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.248032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.248032Z digest=sha256:a325dccae38b36f2c7e57b31dda170fba74e02504997f09a84e8567a11361432

Observation ecbced9a-afae-4de3-ac64-1f0c38fbdb2e · outbound

This paper cites write newline.

Evaluation of Multilingual Image Captioning: How far can we get with CLIP models? write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:58:48.251828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T14:58:48.251828Z digest=sha256:fbb8e651b555e8ebea026f66b7e04c363fc1b2141aa430885b6196127dd22814

Pith citing papers

No inbound Pith citation observations are available.