Pith. sign in

Paper Citation Record · LEDGER

Image-Text Relation Prediction for Multilingual Tweets

As of 23 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2505.05040.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05040 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:18:27.679231Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7fe6c9c8-5327-4cf3-b08f-0a01dfb0a182 · outbound

This paper cites URL: " 'urlintro :=.

Image-Text Relation Prediction for Multilingual Tweets URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.476642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.476642Z digest=sha256:eb6896cf6f9ae57d7e4a06b74324a63bf159d9c2580d66f0a70b260622dfd2b0

Observation 49c3bd12-eb0c-44f2-bed4-e0a6bb34c83d · outbound

This paper cites write newline.

Image-Text Relation Prediction for Multilingual Tweets write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.554419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.554419Z digest=sha256:0cabcc234c3275e38d08e0f0cccf3aa8e5a7416b6e677c7a101dfaece4b754cc

Observation b7000762-ebd1-4c13-bb7c-0c07109c8ff3 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Image-Text Relation Prediction for Multilingual Tweets Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.560387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.560387Z digest=sha256:b0cf2fdb6a8136ebd90445295f726d10f4f466ecf00d32e04e5bb23ea7ecb5ba

Observation 5a2243d6-fe46-475d-9d58-cccc7551ade2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Image-Text Relation Prediction for Multilingual Tweets Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.567913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.567913Z digest=sha256:b72d3c45d179957919dcae6ca06632ceabc4fcec4b0cf4b5c193e9d3c870cca8

Observation 1666e964-2776-43ce-9c3c-97c5f16179e3 · outbound

This paper cites The Llama 3 Herd of Models.

Image-Text Relation Prediction for Multilingual Tweets The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.575335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.575335Z digest=sha256:4fa5f45995b356a26ee9465a6b67d099aaa4af3d3c609813baebf2d6bdcdc22a

Observation 3a24697f-4696-4ce2-b0a4-8f482405c723 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Image-Text Relation Prediction for Multilingual Tweets LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.581476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.581476Z digest=sha256:b91dad4a8f30e2201b94c9db6471a7bcd153f66762aceb216e41b7ce31d8e6f5

Observation e4b338b1-2d6a-4096-b9ec-929d065998e3 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.587122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.587122Z digest=sha256:58f2e7334706c6e6510613c6ddb23816e5fab6adfa37a4ba3e62a04add17c8f0

Observation 67171a4f-ed8f-4116-8667-69037b6919a0 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.592243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.592243Z digest=sha256:ee331f7ab0a2d881b315917cc2febcffe4b0c0c1fa3ed6f92efb921ae5ba123f

Observation 2046a086-0b31-4ddc-b8d0-639c3bc1be74 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.597771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.597771Z digest=sha256:4e7428ae4af99196094fcf1e55eb61e34f90f1df49e65a8d4923540479a27b84

Observation 28e34c86-9794-4d66-8044-9c1fe9ee6924 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.603002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.603002Z digest=sha256:4d0ce338a4c2d388bc0fcfb840bf18b2175609bb4afad44988593633b0fe2c8a

Observation 9a6054fd-98e9-403c-8990-98c4b5068cf4 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:18:27.979232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T23:18:27.635159Z digest=sha256:fad44cda69a0845f45f00ec6996a0968da9332b40352e8b21af9f49b5cfd6d1b

Observation 9b4ce6b8-b676-4064-9a49-128a7ff25869 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 12

Resolution
verified exact
doi, observed 2026-08-15T23:18:27.715860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=arxiv_source observed=2026-08-15T23:18:27.668460Z digest=sha256:e424673b448efd33564f1f5115fc62a62c148a604e86f409bec81f7d7da9284c

Observation 7683990e-3543-49bf-9c93-5e598a5a6fa5 · outbound

This paper cites an unresolved cited work.

Image-Text Relation Prediction for Multilingual Tweets Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.673831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.673831Z digest=sha256:e19177036a173c5f3d3afc59c70ef6aeccf2ae7bee3a1dd788722142c9a67fe9

Observation 37dfd2de-3af0-4c01-b9b1-233381126529 · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

Image-Text Relation Prediction for Multilingual Tweets VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T23:18:27.679231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:18:27.679231Z digest=sha256:ffca9ad43ecc7419faad990c354ae60965c028260ba0757b982f08e7a8234876

Pith citing papers

No inbound Pith citation observations are available.