Pith. sign in

Paper Citation Record · LEDGER

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08710 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:18:51.876640Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af9e9b67-ed51-43f0-a289-8f9970342520 · outbound

This paper cites Learning cnn-lstm architectures for image caption generation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning cnn-lstm architectures for image caption generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.538709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:44.496453Z digest=sha256:32dff92adc76a9ddef7b7d54bc24ca8a3ab436414cbc1a283abb55a018a8a44e

Observation 98920df9-8da1-4c92-b48d-4e976d93e013 · outbound

This paper cites Image captioning through image transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image captioning through image transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.528365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:44.581263Z digest=sha256:873e54c315bf13285c47eae605a0adbebfcd201dddc8e65a334f192566d59bd3

Observation 2b78848c-363e-4dbb-818d-7779191acc0b · outbound

This paper cites Meshed-memory transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meshed-memory transformer for image captioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.517298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:44.763308Z digest=sha256:e035d8c51c0c2f22439d49c0043d2539f745813b3f82852725d988869e963b23

Observation 724a695d-4006-446b-b5d5-c6b7680e669c · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bottom-up and top-down attention for image captioning and visual question answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.505011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:44.883631Z digest=sha256:2d5913eb5da84cc04dfcfeb95831a3ba6b7567f59d196b6969255240c8d1ce87

Observation 3a749b07-904b-437b-a0b3-6be7fa4d1177 · outbound

This paper cites Self- critical sequence training for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Self- critical sequence training for image captioning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.008929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.008929Z digest=sha256:72587dab524588a424c4ebd51d8d2032266e4462f0485a040a85cc2c16430016

Observation 21ca654c-8798-4748-ae07-cac4c4654afa · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.482559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.151174Z digest=sha256:f9e113f52e51e18c03694f7159b4c700653f64a6493c43560cfdf9b33f865245

Observation 6233736e-00d0-45bb-a043-a7cba9c56470 · outbound

This paper cites Microsoft coco: Common objects in context,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Microsoft coco: Common objects in context,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.215989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.215989Z digest=sha256:ef5904e19b4e984d25ab0149f3a126a8a7769754806613afca32aa7ceb8f3777

Observation 8fe3c809-e63e-4f17-ae2a-7ed19de952a0 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bleu: a method for automatic evaluation of machine translation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.325805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.325805Z digest=sha256:1721caeec36e0c64cef5393d6192b451dcaa62286d99faa2c316a90c10b2f631

Observation c7189918-21d6-4020-a0d7-b828d7412d6d · outbound

This paper cites Cider: Consensus- based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Cider: Consensus- based image description evaluation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.401247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.401247Z digest=sha256:2595f150248119c7f2d3b2c45fa4d190838f92adc263de725b956c4b63b94b47

Observation 1a1376e0-81c8-435b-9c09-1c69de0c4497 · outbound

This paper cites Bertscore: Evaluating text generation with bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bertscore: Evaluating text generation with bert,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.451864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.492571Z digest=sha256:03bf4ed80e17dcc288cfadd69de305e393b9c55935baa65e601725c023c417fb

Observation 196d3fa2-6bb0-4897-a325-0db460aa2a3f · outbound

This paper cites What you see is what you read? improv- ing text-image alignment evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training What you see is what you read? improv- ing text-image alignment evaluation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.442462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.585156Z digest=sha256:e3f0c4ca8657605ac765acc7bae74c6906ac4567d7ba9d6c93063132f476d088

Observation 738d87d5-2d8d-40ca-a429-d349426f4e66 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.432955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.683674Z digest=sha256:fb81360066d137a14cda1c664dab20d8e6963aebeaa45b4149f00c202bd42473

Observation 34fd9183-3893-4f64-bb97-884567a1abe8 · outbound

This paper cites Vilbertscore: Evaluating image caption using vision-and-language bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbertscore: Evaluating image caption using vision-and-language bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.421682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.769679Z digest=sha256:c596a407dde78809293a2efa5a9ba35d4abd1057f9a9fd689c506258d69be39d

Observation 16462428-c0fa-474d-a16c-51ad33ce8d5e · outbound

This paper cites Image-text alignment and retrieval using light-weight transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image-text alignment and retrieval using light-weight transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.396031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:45.864014Z digest=sha256:6c7e9a163bacad2de02ec0be8caed7a3c89fdcd3946f0094918389855a3423b7

Observation b86d9278-ae41-4bd3-929a-1f025224a9f0 · outbound

This paper cites UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.987661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.987661Z digest=sha256:4233b782d287cb3e4911a7a8d27852bb22ea5915908000ed79473e981309ceab

Observation b11e06f7-676b-404a-b3d5-0d11a036e1a1 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.382031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.060303Z digest=sha256:3f9995860dd4cbe4b0d155d506f773461450d15fd30b53e2b023ac8f05bc25e5

Observation 523b857d-58f0-4d9e-94a8-81e3b919172a · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.369989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.129387Z digest=sha256:890348095f8f72316b3becd40e889bd65db21d86b9b2f86f38325a4169946c24

Observation 10672b9f-a868-466f-8a6c-b1c479216c38 · outbound

This paper cites Quality estimation for image captions based on large-scale human evaluations,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Quality estimation for image captions based on large-scale human evaluations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.346489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.244803Z digest=sha256:74f6bdab2f076ae8d41bd0c806c3dccd323f58379c30ad8dce7e600f5cb3c28e

Observation 067683fa-cef1-4c83-811c-27ee9be55865 · outbound

This paper cites Revealing the dark secrets of bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Revealing the dark secrets of bert,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.328710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.370999Z digest=sha256:9d9f8fb316101ac7a8fa3d264a72580b865642b472ff9ef1d7384fc25db787c4

Observation 104ac830-39ae-4fed-9f7f-ce993210402d · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Minivit: Compressing vision transformers with weight multiplexing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.318090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.521919Z digest=sha256:bdf7da789b1f2124652dd4eb567f4b152070cbb207bd063d5382993ee26f853b

Observation 8ad54b67-70ec-4ebe-9c0d-6cd9ca908c74 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.607900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.607900Z digest=sha256:18bfb38d7244bcc3be1243473a6ab7cda60f5669ef888f9ae6ea1e575c271077

Observation ee9aa492-63a3-4d64-b84b-55e70a95c154 · outbound

This paper cites Enabling multimodal generation on clip via vision-language knowledge distilla- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Enabling multimodal generation on clip via vision-language knowledge distilla- tion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.291873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:46.699499Z digest=sha256:53546f37529327c79da19433c90d75b1f53eacb2f6b9ab56c9f13290882992a8

Observation 65ec115c-e93a-478d-a9bb-c43a8a8ca597 · outbound

This paper cites CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.816291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.816291Z digest=sha256:d64afcea2edfdb2bdfe1affa0ccc34d5857fedf9bdcdabe6656f4047f439e600

Observation 60ba08d4-fc9b-4506-a040-d906a169dae3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning transferable visual models from natural language supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.902472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.902472Z digest=sha256:6a1a7593ad13d78444497c37eb37e62d15b90ad19e7c71410bb2c2f6a8635909

Observation c0f769db-72fa-4407-b8cd-bcbe1e6ffbd3 · outbound

This paper cites Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.226808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.017461Z digest=sha256:01620a517b9e9801f6180f4f97a045b064e8efc15c3901d607bd90585f445e24

Observation 571e6b59-0e11-4213-bfca-acffcc3b3899 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.182859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.108684Z digest=sha256:1eac890acf7343c9ad7d8489ad3de4f0e1e6dd8cb97a907cbeb76cdd63b61057

Observation c435e904-4efc-468e-90e1-77a151c923e6 · outbound

This paper cites Know more say less: Image captioning based on scene graphs,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Know more say less: Image captioning based on scene graphs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.441150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.214435Z digest=sha256:ce2cbddf152cf0f1778aea334dd0920ca379a080477534659d9f87f2993107c3

Observation 9192a4e7-23b4-4c75-aad3-0e2a41fc051a · outbound

This paper cites High-quality image cap- tioning with fine-grained and semantic-guided visual attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training High-quality image cap- tioning with fine-grained and semantic-guided visual attention,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.142395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.379950Z digest=sha256:b6930bc19278f5e15b74359f7ed022fe52634d976f1601c6f65c5d348203bde2

Observation 45cf7118-9da8-4cb3-b021-9ce38f332a1e · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Multimodal transformer with multi- view visual representation for image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.954306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.450394Z digest=sha256:74858cad36dbb0cfc845bb4366bfbc8037cef9c14eafc9d6a50cbd85fd49dfcd

Observation 9dbb5512-162e-4619-b482-2081258a674a · outbound

This paper cites Compact bidirectional transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Compact bidirectional transformer for image captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.756704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.571278Z digest=sha256:e40ca3ee59d3bc0d18d27d119b8c5d65a2fd054fa23dd926d6fc960c8451d7b7

Observation 58c688ad-e4fc-45fb-a40f-f41e3ad78348 · outbound

This paper cites Task-adaptive attention for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Task-adaptive attention for image captioning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.571741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.661175Z digest=sha256:5c7f2e52ca3e5f13d3de1b3ad77997e32394e0037af6ed9e50ae2f0fdeef6ea3

Observation b8faab56-4325-4121-a7fa-e0fc61873d7c · outbound

This paper cites Textual context-aware dense captioning with diverse words,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Textual context-aware dense captioning with diverse words,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.262369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:47.728248Z digest=sha256:dedae1b5425f6ec44d0e579f00459dedcc2bc2e09293392ec8f9b6c8f367a2de

Observation 46457da1-e0aa-4070-8373-e2adfac91838 · outbound

This paper cites Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.800623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.800623Z digest=sha256:8c760fdad101e597c3c3708856a56178f8a205006d43dd246193c3a653cc8295

Observation a5722b42-588b-466a-be6c-900b74189f3e · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Rouge: A package for automatic evaluation of summaries,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.891704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.891704Z digest=sha256:fc9b34c3e604183c08aff932bc2a5d5763c9a36ddfdad98f9e929c4b96c8f599

Observation b440dc69-0d02-42c1-a27d-3735867b32a8 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spice: Semantic propositional image caption evaluation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.956604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.956604Z digest=sha256:d4402fd189ee1dd47bed41aa297086f68fa4b787efdb073d089641995d67586c

Observation 9266a018-e0eb-40f0-a5ec-d9b3eb341cf1 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.077301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.077301Z digest=sha256:c0b6c2d8de7c4e876a6f5a5c236a7efb4d190414016ead6d3ec1e1c880fdc889

Observation ac99a651-3121-4d31-8668-a42d84296387 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Clipscore: A reference-free evaluation metric for image captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.989883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:48.194780Z digest=sha256:7ca8211541d911de1f403115ae6b52f935009557dbd3b5798495492cd1039f19

Observation 9e56f2de-0a44-46f0-8363-b3366ecad143 · outbound

This paper cites Tiger: Text-to-image grounding for image caption evalua- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tiger: Text-to-image grounding for image caption evalua- tion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.753289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:48.326568Z digest=sha256:b7a5a2b523158a62bc0172c427000cdab81493ecea8045300b9d2b5e9a746e26

Observation 473024b0-5264-48ed-91fa-011520487af3 · outbound

This paper cites Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.505044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:48.435085Z digest=sha256:5a1ef20e9151902073edba50e98b403a82dfbf26eef45033ec2053b000e9cda9

Observation 2bebff7b-075d-4dc3-9477-81a543bc9c20 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Gpt-3: Its nature, scope, limits, and consequences,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.555233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.555233Z digest=sha256:cb2a79afea576ff4e182fca20946efd3f78dab4a7816eb8b8398b504efe5c2ae

Observation 0e0d5cf6-ec6d-4216-9234-d04b0fb41152 · outbound

This paper cites Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.274386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:48.636684Z digest=sha256:16d8d3a19fbd034bf91c5f6741bc5f898bb4e48c7c9efc53155499547e3b68f5

Observation 1b91fbc6-9386-456e-9543-9953ab43f225 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.724702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.724702Z digest=sha256:3e2bcd5bf523fb04a7f7a06019c0d8c557e117dfed1e0386f6e28f638e53c9d6

Observation 2f07502a-6d12-430e-8cbe-bf26a807362b · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.992310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.992310Z digest=sha256:3c37de797573da606e241dbe0f9da472169a047d86fcd0c4968be8237fc10692

Observation e4c820e8-97fb-4d58-8dae-ad870c5dcfc1 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.108576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.108576Z digest=sha256:71ddee4c65e04fa16d2da9aa989d20efa587df0a0f14e70f07e58334da44c017

Observation 671a853a-d350-4185-a7bd-c000fcfee516 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Slip: Self-supervision meets language-image pre-training,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.996841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:49.294938Z digest=sha256:dab478ed092755a5563e9fb1ab71bd696f4cc0172497f7bbba9c14d00a87d0e5

Observation d5a17254-2f11-44e3-be44-f90bd97b8975 · outbound

This paper cites Pruning Filters for Efficient ConvNets.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Pruning Filters for Efficient ConvNets

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.439407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.439407Z digest=sha256:2dde9617a927d08cba3f44c131ebddd95884a929ca9460dc878b8bbbd5ca11e1

Observation ede5327c-922a-4ed8-9555-cb7fb26c0b0c · outbound

This paper cites Vision Transformer Pruning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vision Transformer Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.586224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.586224Z digest=sha256:71485a755c6aa6fd8ff272b0b43c07aea788bd02ad2f9fdf6613cfcfcdf30ce1

Observation df4b184b-157f-4be5-a4cf-3c00dbcc0f2f · outbound

This paper cites Post-training quantization for vision transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Post-training quantization for vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.652256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:49.787638Z digest=sha256:5b096808ebf2dd2127a7bd06e4b008e7ad509c5d625059a818cb336516f74844

Observation 0505ab65-2ef6-496c-98ae-c4ec38866818 · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinyvit: Fast pretraining distillation for small vision transformers,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.444617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:49.905742Z digest=sha256:7472dc414402edb1a432267eb0c1ab842de6b49a9f15e8c1a9c6b854dd46e071

Observation 0b0011c6-d04c-4753-abb2-3a5bdf78d27f · outbound

This paper cites Tinybert: Distilling bert for natural language understanding,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinybert: Distilling bert for natural language understanding,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.259692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:50.076329Z digest=sha256:4a3ec2931d771b55758f3ded1cb552369266c8c900a176e276dd26e0c0726ab0

Observation 2b89c3d1-403f-4093-b9d5-8456c1de0eb0 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.162524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.162524Z digest=sha256:564c5c4aea86d725686b2a13b2411ac57c04c54e9061295bdb1abfeaf236be70

Observation b39e71e5-76e7-4887-a9a7-cdab83f87c18 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Training data-efficient image transformers & distillation through attention,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.303885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.303885Z digest=sha256:6145fc10c5925f3d16b4fed577556ce57ab9cbad82684d6c9dacb4ce6f885c2f

Observation 57e846a7-95fc-4f08-b549-f0006eaa6703 · outbound

This paper cites Attention is all you need,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Attention is all you need,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.460291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.460291Z digest=sha256:92a7d315761dd9702c2a84843c4ab2c37fad0a874ea543a55bbf60f6b183d91b

Observation 507da7cf-40af-4f7e-b50a-a3421316f702 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.608287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.608287Z digest=sha256:7b461c35dedba0529f6aff27f6f5d268735de86f0622a2159b59a1252fa34271

Observation 6b05c5f1-1756-47bf-91a4-56717f2c1ae2 · outbound

This paper cites Bag of tricks for image classification with convolutional neural networks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bag of tricks for image classification with convolutional neural networks,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.104734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:50.761420Z digest=sha256:7e9945925c41c6e3ad45a26855090dd3bf3606523ecb8255015ca797312156ae

Observation 4af3bb92-6134-416d-b4aa-62c6937a5550 · outbound

This paper cites Making convolutional networks shift-invariant again,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Making convolutional networks shift-invariant again,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.800069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:50.886873Z digest=sha256:8d31c5d8fcce59d10185ba112ec89d8aef0b9265f78f7bd068bd4b074caecede

Observation 20d691ff-b298-4771-935b-70f04e205385 · outbound

This paper cites Discriminability objective for training descriptive captions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Discriminability objective for training descriptive captions,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.570868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:50.986021Z digest=sha256:6d1692d26d1ee15099a1c86b2e2a1f0a1a0cf436234ce0c45ece3cdf476d2f32

Observation 07f51bdc-c850-49b2-81e8-fb98b53f422f · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ImageNet Large Scale Visual Recognition Challenge,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.056134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.056134Z digest=sha256:2632aa9e907693c2f220b9b4da0e7d0cc313cf8f9c4de0e899e2a4cb888d3bc7

Observation 2ba5db4d-d95c-41f0-a492-b0611a25abff · outbound

This paper cites Framing image description as a ranking task: Data, models and evaluation metrics,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Framing image description as a ranking task: Data, models and evaluation metrics,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.266631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.171017Z digest=sha256:2f0c38d27edd1a2a6d9fe81c9c7c691db8bbb86eec020320327a88b3b1b6ff32

Observation fa8cb4f5-4a9b-4e46-b2d8-af3aa3151a8b · outbound

This paper cites From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.241514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.241514Z digest=sha256:09c0cea12c589710f8485365462eab1940e86bbeb28813168e27c3f037873c85

Observation b5ca7f10-23d0-4445-9220-275089ddd46e · outbound

This paper cites Consensus-based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Consensus-based image description evaluation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.042948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.305407Z digest=sha256:4a53e286e4ee62e4178137fe706e7a4c894e1cf429c9646076b800df931722f1

Observation 0433243b-fd85-4fcc-af24-11e652ff5144 · outbound

This paper cites Foil it! find one mismatch between image and language caption,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Foil it! find one mismatch between image and language caption,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.700175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.395791Z digest=sha256:4fd0e706af43c24332d0c43c144d68900d020903ae9ff2b74711605e1206253f

Observation 6705d649-e7d7-4e3f-9634-a4129066491c · outbound

This paper cites Deep visual-semantic alignments for gen- erating image descriptions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Deep visual-semantic alignments for gen- erating image descriptions,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.538573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.477940Z digest=sha256:2967f66dbf69c99464b29eaf4fc658918cc9783033ca0ffc74742c61579fe468

Observation 24f8f060-4a72-4d69-a4fb-4b7d196df7b4 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vinvl: Revisiting visual representations in vision-language models,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.392262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.554693Z digest=sha256:14d2d1fd88d4eca0cded74c5f39fdb22eba134bb134b0e38dc2139b49c51229b

Observation 8dd056fb-e566-42b9-a6c3-2a426bbf0ca9 · outbound

This paper cites Adam: A method for stochastic optimization,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Adam: A method for stochastic optimization,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.258827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.640789Z digest=sha256:141fb31b44fcc5574f4de98176ca130020b59b66fa26eb7ed4b41377055bb293

Observation fc4aa0b8-2168-462a-8dd9-58111a0fc563 · outbound

This paper cites Improving image captioning evaluation by considering inter references variance,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Improving image captioning evaluation by considering inter references variance,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.129394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-08-06T18:18:51.705906Z digest=sha256:bf147a91bbcf968425d422ccf490da5cf3abd5467296b0062d299bf89354a787

Observation 380b0f08-6362-4ce7-8203-5776db952ae8 · outbound

This paper cites Concrete Problems in AI Safety.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Concrete Problems in AI Safety

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.809319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.809319Z digest=sha256:7dca518f0feb04a69d890c843ce247115e10eca1d6728dbe64afad433a46edd9

Observation 61a01979-bf04-494a-9ff6-eac046908ce5 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.876640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.876640Z digest=sha256:c83bc5e7c2a0d84b9fa8f126e1e44c349ee835f64fd45250607f87ecb152a43e

Pith citing papers

No inbound Pith citation observations are available.