Pith. sign in

Paper Citation Record · LEDGER

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval

As of 23 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2608.05260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.05260 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T17:00:33.291636Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved8
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1a5fdd20-2aab-47c5-853e-d03d0343c3c5 · outbound

This paper cites In: Eu- ropean Conference on Computer Vision (ECCV).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Eu- ropean Conference on Computer Vision (ECCV)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.839636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.183793Z digest=sha256:806c032319bdac4e4beed3c56e208095d8d93d20fbbe0dd0b055cee120a7637d

Observation c84c8f52-7611-4bc3-9519-a9dba28dc1ae · outbound

This paper cites Advances in Neural Information Processing Systems (NeurIPS) 36, 35544–35575 (2023).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval Advances in Neural Information Processing Systems (NeurIPS) 36, 35544–35575 (2023)

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.824934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.189678Z digest=sha256:ea30d7a2c6d79dad3c36381ee62dcd08908175d28a8c47baec677a8bf512f144

Observation fca33485-8856-4111-b7c8-04c459bc969d · outbound

This paper cites The Llama 3 Herd of Models.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval The Llama 3 Herd of Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.194553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.194553Z digest=sha256:1337df4b3a5f473dc39dca5f95b2b2f3c51a32922569b43fd309e986f4485853

Observation a441bcaa-489b-4f15-8d42-50f5c84fc98a · outbound

This paper cites In: International Conference on Machine Learning (ICML).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Conference on Machine Learning (ICML)

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.808782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.199771Z digest=sha256:db0f27507238baf5e119c51b6f40064c6f6dc6df3fef7e640914f07964346763

Observation 747a5b31-3f90-40e8-a59a-8eb38c60963b · outbound

This paper cites In: European Conference on Computer Vision (ECCV).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: European Conference on Computer Vision (ECCV)

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.793948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.205035Z digest=sha256:ce09573441017ca9b158c68271f8fa3892ba3b3969ddfb7d0fade1561b77a761

Observation 772ada05-14a1-4ea9-a61f-86d3b92fd15b · outbound

This paper cites In: Proceedings of the International Conference on Machine Learning (ICML) (2024).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceedings of the International Conference on Machine Learning (ICML) (2024)

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.779205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.210160Z digest=sha256:a370ddfd9f8e4407f4a8463feb93dde2d8bc917ad2c8ad66bfb8a6010b45ca07

Observation 1803b72a-b9a2-4071-a272-928f23657fa6 · outbound

This paper cites In: International Con- ference on Machine Learning (ICML).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Con- ference on Machine Learning (ICML)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.763972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.215562Z digest=sha256:cc5515a1194ef446a8cc248542c29e9fd788ee3bc846e4788503d34e0f459ade

Observation b0412ad6-88e7-402e-86f7-3256a986a781 · outbound

This paper cites In: European Conference on Computer Vision (ECCV).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: European Conference on Computer Vision (ECCV)

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.748984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.220176Z digest=sha256:ef3c242a719a4eb67b95cfa44123a58ffda61f1ccf9277a62f4d87b95b701910

Observation 3146cad1-0bb5-4ab9-921e-c03d3e4424cf · outbound

This paper cites In: International Conference on Learning Representations (ICLR) (2019).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Conference on Learning Representations (ICLR) (2019)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.224619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.224619Z digest=sha256:d03e1fb35a2bc1601dc8e87f5e36a46471e72ead6bc3898ae2be2f88619c8fe0

Observation 82168643-d68e-4d00-8a8b-0d19d2f33991 · outbound

This paper cites In: ECCV (2024).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: ECCV (2024)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.724666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.228860Z digest=sha256:802f23cec8887946330beb40308eda9b37531d98010fbaf22e7ab23e7da3811d

Observation 77424ed7-2477-4219-b5b0-588be0c1c8d4 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval Representation Learning with Contrastive Predictive Coding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.233154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.233154Z digest=sha256:f8f662bc29db118d4ec9cd4c0ad93fd8a14382ab7ad1a0c2a67c18d8f87a580c

Observation d7054a31-e0ff-4599-a479-8cdf76b95375 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.709482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.237318Z digest=sha256:d080b0ed2f0fd7101ba305e2dfa83945fc5a9d8163fa9f3f35b7b0a3aeb09d15

Observation 704290da-8d07-4d16-9340-84be04ed005a · outbound

This paper cites In: International Conference on Machine Learning (ICML).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Conference on Machine Learning (ICML)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.241628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.241628Z digest=sha256:71f2ec3793307c755f2e3b2f30b266f7ba2cc5686687111cb4aa2814ee4c6aed

Observation 84de665c-24d5-4ef3-8546-ff4ec788a225 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.245583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.245583Z digest=sha256:72569a5b1b008672520ae3a00b2422d845818770ffebf795b8474acc1daf60a9

Observation 08688c6e-4711-48dc-aeb4-d5b522261793 · outbound

This paper cites In: International Conference on Learning Representations (ICLR) (2021) 16 M.Ghazanfari et al.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Conference on Learning Representations (ICLR) (2021) 16 M.Ghazanfari et al

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.685899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.249586Z digest=sha256:b5d85855bf0eecda0d5dc3a1610ff51fb5ee2442821641c6bf5594c6ae26c0cb

Observation 7ee4f2b7-6c01-4a99-813e-bbae9f47fceb · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.253337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.253337Z digest=sha256:9e2cc2dac91fb522a1c7d881c2b3eaf590639465fb6e7d43f6f0cf912cf14e04

Observation e98ffb44-07fb-41d4-8f9b-c2ba80873e36 · outbound

This paper cites In: Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceedings of the 3rd Workshop on Advances in Language and Vision Research (ALVR)

Reference 17

Resolution
malformed identifier
doi_truncated, observed 2026-08-08T17:00:33.330630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.257190Z digest=sha256:54da40eac8f87c6237481996cab7d5e039997561a312f021d26c8bd7ab10d926

Observation dee9fbdb-d40e-43a4-a8eb-ab9774a627c4 · outbound

This paper cites In: Proceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceed- ings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.659977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.261074Z digest=sha256:42c7f240896efd4e92244514468b9c17f8155fdf5899397d956160a5e373de5b

Observation 83df07ab-95f4-4070-aed7-f1861d75ff57 · outbound

This paper cites In: International Con- ference on Learning Representations (ICLR) Workshop (2014).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: International Con- ference on Learning Representations (ICLR) Workshop (2014)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.643309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.265260Z digest=sha256:2386a727585521f95d491ff534ff24b3e1b1bcfb359807577cb043f615607e12

Observation 9ca25fc7-b458-42fd-ba49-27427a49dec9 · outbound

This paper cites In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: 2025 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.269361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.269361Z digest=sha256:bb0189b5440cca6c28efde0bd945b6cadd5fd7d5b19c37682eb717de8cd2c8ca

Observation 996660f5-0810-40a4-9a7c-492b160e3f47 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T17:00:33.273928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:00:33.273928Z digest=sha256:8541a626541313f3f0abafc0117fe9bb892abd776736afad02d02643dbca5d2c

Observation 280b3164-6e06-4b20-a3d7-571b10e6f597 · outbound

This paper cites Transactions of the Association for Computational Linguistics (ACL)2, 67–78 (2014).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval Transactions of the Association for Computational Linguistics (ACL)2, 67–78 (2014)

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.627951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.278526Z digest=sha256:6829daf49bda40e71eaae5266ae0c8194dfd6259163da3160830764619cd3c41

Observation 5b3ea0d7-8b06-47d6-b2f2-809ae35e2167 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.612458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.282603Z digest=sha256:3aec131fe44ca52f1713611a587906c439559fdc9560a8820c4115e519f59193

Observation 785dc7fe-8379-4fc9-8bba-037e1515c7a9 · outbound

This paper cites In: European Conference on Computer Vision (ECCV).

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval In: European Conference on Computer Vision (ECCV)

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T17:00:33.597247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.286962Z digest=sha256:c6deaffd904cce60f5916ac693fd610bf5a08f263700f9a35968daaa60915c88

Observation 66e68059-776e-4c4d-af48-b09ab75109e9 · outbound

This paper cites TRUTH / IS A BREATH AWAY / THE VICTORY.

A Paragraph is Worth a Thousand Captions: Rethinking Text Supervision for Vision-Language Retrieval TRUTH / IS A BREATH AWAY / THE VICTORY

Reference 25

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T17:00:33.581371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-08T17:00:33.291636Z digest=sha256:7a025332b44223443446b0164847d65bd1da67a275bd3ddcf6a8e0c4d4b9e258

Pith citing papers

No inbound Pith citation observations are available.