Pith. sign in

Paper Citation Record · LEDGER

Multilingual Training and Evaluation Resources for Vision-Language Models

As of 31 July 2026, this Paper Citation Record lists 44 of 44 outbound references and 0 inbound Pith citation observations for arXiv:2604.18347.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.18347 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-05T12:25:00.027624Z

measured 44 of 44 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-07-31T06:34:12.847434+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact6
  • verified fuzzy28
  • unresolved2
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a7fbb15e-dfac-4b9f-af9a-0d9ee1cb9692 · outbound

This paper cites GPT-4 Technical Report.

Multilingual Training and Evaluation Resources for Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.863814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:20d72cd7d3330c9616e08f60c98927120a4630b8a65660404c53d8e8fb7cb877

Observation 4ddf72b3-be2c-4154-b12e-3fc7a1ffd9dd · outbound

This paper cites In: NeurIPS (2022).

Multilingual Training and Evaluation Resources for Vision-Language Models In: NeurIPS (2022)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.084671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:09949659cd957e7e7c61f6af3f189e4dc16b06fc0417b8869396fdc7bb16264e

Observation 9d3da8ac-99af-416d-8311-6487b6343f01 · outbound

This paper cites In: ICCV (2015).

Multilingual Training and Evaluation Resources for Vision-Language Models In: ICCV (2015)

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.065482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:af59d6d14f40664e4acc1c5a3cfa479af6a3f6f37a99bfb1d334805f01eb3e9f

Observation eeeac37f-c9d2-4509-9c68-e29fb4d96c49 · outbound

This paper cites Large image datasets: A pyrrhic win for computer vision?.

Multilingual Training and Evaluation Resources for Vision-Language Models Large image datasets: A pyrrhic win for computer vision?

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.838782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:52b7306457b347b395ed5bc8ef5fdc74bfc84cd8c147738027abd099bd538ef4

Observation 9e3d1856-9a88-43b6-a4e0-a7055de32dbc · outbound

This paper cites FAccT (2021).

Multilingual Training and Evaluation Resources for Vision-Language Models FAccT (2021)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.060241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:b28f7b407b4b3acb97559d596b72dea82520d62f4daacb970dd8c6b4aafa473e

Observation c1bb11a9-ec64-4e68-b0c1-759f78444bee · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Multilingual Training and Evaluation Resources for Vision-Language Models Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.865337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:f6c2e2c19b0bca71c873e2fc2b7ed8551a71ebc049b098911fc43467c33b2c3e

Observation 9211162a-1a61-40cb-88cc-61186fcf42aa · outbound

This paper cites In: ACL (2020).

Multilingual Training and Evaluation Resources for Vision-Language Models In: ACL (2020)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.071541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:6a423f037af94f0a9b620e121fa024b810a9d329ec3b0064b62223c42fba01c4

Observation 27315e54-13b5-49ba-a186-1163096095af · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

Multilingual Training and Evaluation Resources for Vision-Language Models Unsupervised Cross-lingual Representation Learning at Scale

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.858791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:88e92214aa3f8c0c0925a34ff36d58b5a8f9f8e77aafae1ba86fc1ff4289dee2

Observation 7d957ac8-78dc-4a30-8014-452a1a00af34 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025).

Multilingual Training and Evaluation Resources for Vision-Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2025)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.075442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:8a8dd92576f555fe0b562379562e821b988a69926338ec7e711d3cca4d56dc35

Observation f4148607-ac8e-42da-aecf-855cf5ebc93c · outbound

This paper cites In: Proceedings of the 32nd ACM International Conference on Multimedia.

Multilingual Training and Evaluation Resources for Vision-Language Models In: Proceedings of the 32nd ACM International Conference on Multimedia

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.087228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:3cec6a97b62ca072b3fade1828bb5412660ff023940c588607587389f413677f

Observation 685df082-9469-422b-aebf-c72460905ac7 · outbound

This paper cites In: VL@ACL (2016).

Multilingual Training and Evaluation Resources for Vision-Language Models In: VL@ACL (2016)

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.075239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:7e3538d780f0625bb551a25c8670e55289f056b67a5bad8eb5630f3c55e88292

Observation 78306f1d-bec3-4c61-a8a4-b7a8f1880235 · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track (2023).

Multilingual Training and Evaluation Resources for Vision-Language Models In: Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track (2023)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.032667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:e80263a6f2e35945da1f6f1fbf00b0b9b4e10cb76e5cf9aef638daa92e46347c

Observation 8180e686-f5c9-49d7-a093-2c32cc399342 · outbound

This paper cites HuggingFace model card (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace model card (2024)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.081130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:1f6ed06ea3e276cde856ea181094bea48f548da630e6d510e4e45c0568f2a809

Observation 90e626c7-1cca-4fb3-896f-84e1c628444a · outbound

This paper cites International Journal of Computer Vision127(4), 398–414 (2019).

Multilingual Training and Evaluation Resources for Vision-Language Models International Journal of Computer Vision127(4), 398–414 (2019)

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.063505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:6f21ab0d22b4507bfffecc2b6d69530b912cbc762bcba7a40b7ec37f687e8eb5

Observation 350d19db-9371-4ad9-85b8-7b0cbd0dc311 · outbound

This paper cites The Llama 3 Herd of Models.

Multilingual Training and Evaluation Resources for Vision-Language Models The Llama 3 Herd of Models

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.884326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:38654bda229c8a3a6bd7381973d70a1e1cf6d2ab79d510e77256275e64efafe7

Observation de0d56ca-f179-4383-a139-194fa10790b3 · outbound

This paper cites British Journal of Mathematical and Statistical Psychology61(1), 29–48 (2008).https://doi.org/10.1348/000711006X126600.

Multilingual Training and Evaluation Resources for Vision-Language Models British Journal of Mathematical and Statistical Psychology61(1), 29–48 (2008).https://doi.org/10.1348/000711006X126600

Reference 16

Resolution
verified exact
doi, observed 2026-07-05T12:30:59.740528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:dc5a81a0a102a23505bc69c53c8f9fe0a61a1953a86e03d2be77b4cb62afee50

Observation 5a1a2968-a008-47d7-ace4-b3c050293a5e · outbound

This paper cites HuggingFace dataset (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.044550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:4e2a3c2e128e2d510b41bf90ad170108915ec321ebbff8729389dcb1ebd1a620

Observation 168b4f8d-16d6-4e14-8899-0ac10eb7b417 · outbound

This paper cites HuggingFace dataset (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.069378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:7f1a0b43137d6c4741a466c394a797e52833640043626aaa3fe5f51896389fc1

Observation 4cc32fc1-b8a9-4a81-9ec4-8fcf327f3c41 · outbound

This paper cites HuggingFace dataset (2024) 16 Baiamonte et al.

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024) 16 Baiamonte et al

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.073295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:37fd2ca3690e429758b29abb4f99c2e7d600cca3129e83145de475dd40861163

Observation 34d3d6d9-f05c-4a84-b9d1-87f25b5b1fdc · outbound

This paper cites HuggingFace model card (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace model card (2024)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.076886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:d6844816847c872440c426aea97ea1fa7a547901ca41b66f360fd61e53028e62

Observation bf54d0dc-c9f5-4b41-8114-f3b785d00f4f · outbound

This paper cites In: ECCV (2016).

Multilingual Training and Evaluation Resources for Vision-Language Models In: ECCV (2016)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.051741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:8143ebe10ca77c06314398bb9cdac1c784b48434838871ea75d8f7778a8a66f9

Observation ef74b3e1-7dd6-4f0a-9639-d19b6552aabb · outbound

This paper cites an unresolved cited work.

Multilingual Training and Evaluation Resources for Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-05T12:31:00.082098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:4ee8e5a3119eae89b7750b49b5252333c2201cba24914f3236236410e2018d85

Observation 6d6ce162-132d-45d7-be0c-2a148a666db9 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

Multilingual Training and Evaluation Resources for Vision-Language Models BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 23

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.881443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:df32fc092caabddf8ca7f8c7deec6e23e0173c20df753ec0f174a408e833a994

Observation 5495ec88-a050-4ba3-8af8-780d96ef3d3d · outbound

This paper cites Integral Probability Metrics on submanifolds: interpolation inequalities and optimal inference.

Multilingual Training and Evaluation Resources for Vision-Language Models Integral Probability Metrics on submanifolds: interpolation inequalities and optimal inference

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.880961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:9da90c93dd02b1caf630518e90351e92d51b6012afd6a1a3e8927449abda051a

Observation 4e0c5609-8ac1-4c18-b1b7-02a08f3666c5 · outbound

This paper cites In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2023).

Multilingual Training and Evaluation Resources for Vision-Language Models In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP) (2023)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.083982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:4624d73d65e235316465443a48844b4df942e7d757f10fe4930f721199888a9a

Observation 3f063a40-dd45-4f66-87bc-f8abed9ab106 · outbound

This paper cites an unresolved cited work.

Multilingual Training and Evaluation Resources for Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-07-05T12:31:00.085575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:c004c3d73b343b8af1a2fe3777f4467e0f1241b4530591f900d58919618ba40c

Observation 4725a56c-d546-461a-b9b3-0ae4a5e9e249 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Multilingual Training and Evaluation Resources for Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.887960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:bf80484cb09a60199ca5c3122af3c76f2a0a7557a10591124478cad1422291df

Observation b0cc5f1d-1b87-4380-8b07-ac3eb80d8798 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (2023).

Multilingual Training and Evaluation Resources for Vision-Language Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (2023)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.077211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:c15834b6ad8a0386f39e49797b0d3685bb7e4f665e754c6aee0450b9bab3d203

Observation cff13f57-7794-4acb-b396-87c4048c0888 · outbound

This paper cites HuggingFace dataset (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.048163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:2eb74fb22a40c4a7811678e314b274e25c16e032b965746a3f06f5abf3e5c008

Observation 728a7eb6-24d9-4b42-beba-61561da49cb6 · outbound

This paper cites HuggingFace dataset (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024)

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.083048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:93f86bda0500020517bbe024a7cf56abe5de7c5318c96661b10442ad67de4994

Observation b8ccff5b-9340-46ac-8d6b-edc1a36b9318 · outbound

This paper cites In: Advances in Neural Information Processing Systems (NeurIPS) (2022).

Multilingual Training and Evaluation Resources for Vision-Language Models In: Advances in Neural Information Processing Systems (NeurIPS) (2022)

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.088132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:b89d31d310bf6f70d3034bf0cda4110647a15815d88d4b39725a611b0b9a25ef

Observation b41e94bc-ea87-4713-a697-031e4d4d26a5 · outbound

This paper cites In: ACL (2022).

Multilingual Training and Evaluation Resources for Vision-Language Models In: ACL (2022)

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.086391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:40c5657b06ee9d9efec774ba5563656683db1f56c8dd979d563231e1f717cd8a

Observation 281fcdb0-1b71-4006-a352-aa86eb4a1f98 · outbound

This paper cites In: WACV (2021).

Multilingual Training and Evaluation Resources for Vision-Language Models In: WACV (2021)

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.063710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:687e727df7ff8135979911407e95eedf341c8078eb99c76414ceeeceda10e3f1

Observation 54f12ae7-319d-4c9e-99a3-5ba74c0d8715 · outbound

This paper cites HuggingFace dataset (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2024)

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.048374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:3811405df353a2a4a2d2420d79262eb61cc49fb4a804fef7ce19cbc36e7d908b

Observation 1c680fd4-5e7c-4f74-8b08-0514715c67ed · outbound

This paper cites Patterns (2021).

Multilingual Training and Evaluation Resources for Vision-Language Models Patterns (2021)

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.061762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:62d372a9bd24718201862ea3956f98dd5de3ec59fd970c741948bf1fc55c72de

Observation 8eb6b9d8-eb7a-4af7-b9f1-7072956b1f6c · outbound

This paper cites HuggingFace model card (2024).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace model card (2024)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.036592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:82936e3116665c8c4fd88566b2cf0e724d2cc1154b629dc8bcc5cad4a34879c2

Observation 118471df-10b7-457a-a494-9fbd9763093a · outbound

This paper cites In: International Conference on Machine Learning (2021).

Multilingual Training and Evaluation Resources for Vision-Language Models In: International Conference on Machine Learning (2021)

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.073097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:fcdfdbe388e490403fb5d5e720f558038b7c4e4e7c02955f340635061508d709

Observation 13c96463-94c8-4a02-951e-b88da872e456 · outbound

This paper cites In: CVPR (2019).

Multilingual Training and Evaluation Resources for Vision-Language Models In: CVPR (2019)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.071408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:c8927fd926ecea4e454de052b92f7a5bfbe1eb5f82bef380e64c621469fffd91

Observation d132e84b-6c28-4475-8e41-eb619ea4a887 · outbound

This paper cites MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering.

Multilingual Training and Evaluation Resources for Vision-Language Models MTVQA: Benchmarking Multilingual Text-Centric Visual Question Answering

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.869841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:360fd6f4968090ca44b4e9bbce4a19667d30ea483abd7107775130814cf941e7

Observation 3565b1f1-91bc-4984-b369-9bca879a2396 · outbound

This paper cites Qwen3 Technical Report.

Multilingual Training and Evaluation Resources for Vision-Language Models Qwen3 Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.877464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:cc98d0b0deb33572e98579b7ab260a6738a492565dd3f7075fca5ee3c24004c6

Observation e6418c5a-40eb-46f5-b371-bdc14dbf260d · outbound

This paper cites HuggingFace dataset (2023).

Multilingual Training and Evaluation Resources for Vision-Language Models HuggingFace dataset (2023)

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-05T12:31:00.067719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:ee592a30ba6901625ad890026598d42cb0d5ddedee6fd3225d7e8290193a0bbf

Observation 143edbfc-d1ba-4bfd-be0c-1568758d717c · outbound

This paper cites Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation.

Multilingual Training and Evaluation Resources for Vision-Language Models Scaling Text-Rich Image Understanding via Code-Guided Synthetic Multimodal Data Generation

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T12:30:59.860761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:821675fd438dd6bd8df15cfd2982462c0e9cc19214b2921a933d8a5ee76d1fe5

Observation 0584df97-beb1-494f-9d5b-60198d09e5d4 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

Multilingual Training and Evaluation Resources for Vision-Language Models BERTScore: Evaluating Text Generation with BERT

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-05T12:30:59.889621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:1a0f6d69d5055cfd6e28508a5068cf0a196387d199abdba484970a9395677daa

Observation bd2d77e1-ac2c-4147-b434-d7a6b0bde66c · outbound

This paper cites Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs.

Multilingual Training and Evaluation Resources for Vision-Language Models Performance Portable Monte Carlo Particle Transport on Intel, NVIDIA, and AMD GPUs

Reference 44

Resolution
malformed identifier
local_arxiv, observed 2026-07-05T12:30:59.884018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-07-31T06:34:12.847434+00:00.

source=pdf_text observed=2026-07-05T12:25:00.027624Z digest=sha256:3aff4790848843eaaa0592cc832f745482b23fd9c37ee1ad799c64b7962175d8

Pith citing papers

No inbound Pith citation observations are available.