Pith. sign in

Paper Citation Record · LEDGER

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training

As of 7 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2507.08710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08710 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:18:51.876640Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact0
  • verified fuzzy42
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation af9e9b67-ed51-43f0-a289-8f9970342520 · outbound

This paper cites Learning cnn-lstm architectures for image caption generation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning cnn-lstm architectures for image caption generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.538709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:44.496453Z digest=sha256:3c2c47900f5e327f3649ffa406c3f04a365ca153acaccb725d45be44bde4996a

Observation 98920df9-8da1-4c92-b48d-4e976d93e013 · outbound

This paper cites Image captioning through image transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image captioning through image transformer,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.528365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:44.581263Z digest=sha256:9ba8172d6ffa75045723af6ed9f5a504b07b238519776cc37b63786a75936837

Observation 2b78848c-363e-4dbb-818d-7779191acc0b · outbound

This paper cites Meshed-memory transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meshed-memory transformer for image captioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.517298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:44.763308Z digest=sha256:c5982329609edcb6b1709e31db4376f0eafd9ba4530ca7d4cef65ef6cac561ad

Observation 724a695d-4006-446b-b5d5-c6b7680e669c · outbound

This paper cites Bottom-up and top-down attention for image captioning and visual question answering,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bottom-up and top-down attention for image captioning and visual question answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.505011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:44.883631Z digest=sha256:4d17e461204ced1ad86fd59bc9cdfd57419c710e5dbcb84bf4b8aa67dc5a4113

Observation 3a749b07-904b-437b-a0b3-6be7fa4d1177 · outbound

This paper cites Self- critical sequence training for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Self- critical sequence training for image captioning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.008929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.008929Z digest=sha256:8725b587aa1332449afd8799c74677ffa693d738563554c3a48f64c49e050c9a

Observation 21ca654c-8798-4748-ae07-cac4c4654afa · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image cap- tioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.482559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.151174Z digest=sha256:45c34aa1ed37f37a77a330e014414fad75d72558980689613008986cdd9abeea

Observation 6233736e-00d0-45bb-a043-a7cba9c56470 · outbound

This paper cites Microsoft coco: Common objects in context,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Microsoft coco: Common objects in context,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.215989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.215989Z digest=sha256:4b1027f045b43c5bff08e04cab24f71defe7e88772d5e00494e4dc78c9181c3b

Observation 8fe3c809-e63e-4f17-ae2a-7ed19de952a0 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bleu: a method for automatic evaluation of machine translation,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.325805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.325805Z digest=sha256:ffa507c2bb727cf82026504aecc9811e1a334ffda82dfe3dfb0f67bd17244dcc

Observation c7189918-21d6-4020-a0d7-b828d7412d6d · outbound

This paper cites Cider: Consensus- based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Cider: Consensus- based image description evaluation,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.401247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.401247Z digest=sha256:6a5e301e392cce4e2f7ea9f4abdccdd81929124770322c428bf2f30ae7f709d5

Observation 1a1376e0-81c8-435b-9c09-1c69de0c4497 · outbound

This paper cites Bertscore: Evaluating text generation with bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bertscore: Evaluating text generation with bert,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.451864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.492571Z digest=sha256:700020a05b1de388dbf1f70c532ddf853a558ac3a1bce2fe511b30b0ff380468

Observation 196d3fa2-6bb0-4897-a325-0db460aa2a3f · outbound

This paper cites What you see is what you read? improv- ing text-image alignment evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training What you see is what you read? improv- ing text-image alignment evaluation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.442462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.585156Z digest=sha256:bf89bb7bb7b4eb1057741a644ee55023f307687b86ef6b05190d5ce2f103ce9d

Observation 738d87d5-2d8d-40ca-a429-d349426f4e66 · outbound

This paper cites Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Fast, accurate, and lightweight memory-enhanced embedding learning framework for image-text retrieval,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.432955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.683674Z digest=sha256:6a17bf1b1681f6d0fd8028450f65bf444258d84e094d3f4d2796fbf64a9aa67f

Observation 34fd9183-3893-4f64-bb97-884567a1abe8 · outbound

This paper cites Vilbertscore: Evaluating image caption using vision-and-language bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbertscore: Evaluating image caption using vision-and-language bert,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.421682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.769679Z digest=sha256:5a2551a820d4c01f73f166016fb7e753a638cd9736f9277b62395b4fbfc7aaf0

Observation 16462428-c0fa-474d-a16c-51ad33ce8d5e · outbound

This paper cites Image-text alignment and retrieval using light-weight transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image-text alignment and retrieval using light-weight transformer,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.396031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:45.864014Z digest=sha256:b0b1254ed8357026400c60b83875cbc0c3401465b100aba12434fbf205b99e38

Observation b86d9278-ae41-4bd3-929a-1f025224a9f0 · outbound

This paper cites UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training UMIC: An Unreferenced Metric for Image Captioning via Contrastive Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:45.987661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:45.987661Z digest=sha256:362cdefd7f659715b7b5bcf7a370a0331840295716a14561b9fd3519426dc271

Observation b11e06f7-676b-404a-b3d5-0d11a036e1a1 · outbound

This paper cites Less is more: Clipbert for video-and-language learning via sparse sampling,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Less is more: Clipbert for video-and-language learning via sparse sampling,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.382031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.060303Z digest=sha256:b764f4fbcb0d05b3d199746af9b35fa41aa99cbbe1c6435bf90c1215d33ff140

Observation 523b857d-58f0-4d9e-94a8-81e3b919172a · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.369989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.129387Z digest=sha256:4623044d3322e13272dace88eca83a7055aa20460cf63879f4b7b3e60872bff9

Observation 10672b9f-a868-466f-8a6c-b1c479216c38 · outbound

This paper cites Quality estimation for image captions based on large-scale human evaluations,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Quality estimation for image captions based on large-scale human evaluations,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.346489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.244803Z digest=sha256:ee27c3fc462bae63b52d67a913eca758c388f6db57c04a40be851c8afa19aec9

Observation 067683fa-cef1-4c83-811c-27ee9be55865 · outbound

This paper cites Revealing the dark secrets of bert,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Revealing the dark secrets of bert,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.328710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.370999Z digest=sha256:41bfd60b8055b49387f16a6f707e1d52e5effce4b59dc8bc016eb65122916947

Observation 104ac830-39ae-4fed-9f7f-ce993210402d · outbound

This paper cites Minivit: Compressing vision transformers with weight multiplexing,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Minivit: Compressing vision transformers with weight multiplexing,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.318090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.521919Z digest=sha256:624b36b96ce4cdba7ba0f45b99581110ea7360be0b7c5133d2d366283a5dd2dd

Observation 8ad54b67-70ec-4ebe-9c0d-6cd9ca908c74 · outbound

This paper cites ALBERT: A Lite BERT for Self-supervised Learning of Language Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ALBERT: A Lite BERT for Self-supervised Learning of Language Representations

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.607900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.607900Z digest=sha256:73d19d3757fa228952cf0e16d941fc9bde4afcc8fa60f880ba7e355873d95059

Observation ee9aa492-63a3-4d64-b84b-55e70a95c154 · outbound

This paper cites Enabling multimodal generation on clip via vision-language knowledge distilla- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Enabling multimodal generation on clip via vision-language knowledge distilla- tion,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.291873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:46.699499Z digest=sha256:3acebd8caf6b5d8126701b666b299811f6a0349683af0c8d66b74b14a665a8d1

Observation 65ec115c-e93a-478d-a9bb-c43a8a8ca597 · outbound

This paper cites CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training CLIP-TD: CLIP Targeted Distillation for Vision-Language Tasks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.816291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.816291Z digest=sha256:f793e52e35596bd64cb624165a26c08af61c2f52168e8142ef47e6540e9fafd6

Observation 60ba08d4-fc9b-4506-a040-d906a169dae3 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Learning transferable visual models from natural language supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:46.902472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:46.902472Z digest=sha256:c33dcef6a365b2fe630110dd8922084d4bdc472f78843217bb01c87c60650497

Observation c0f769db-72fa-4407-b8cd-bcbe1e6ffbd3 · outbound

This paper cites Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Sca-cnn: Spatial and channel-wise attention in convolutional networks for image captioning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.226808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.017461Z digest=sha256:9eaf44ad8445e06d9714adf33c3e1f3268a70c83cc58d1572e6017236c692c70

Observation 571e6b59-0e11-4213-bfca-acffcc3b3899 · outbound

This paper cites Knowing when to look: Adaptive attention via a visual sentinel for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Knowing when to look: Adaptive attention via a visual sentinel for image captioning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:19:58.182859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.108684Z digest=sha256:14ab3aae1e87d141078102002c9175aeddbd9e48bd619adecdfeaa57c45ece02

Observation c435e904-4efc-468e-90e1-77a151c923e6 · outbound

This paper cites Know more say less: Image captioning based on scene graphs,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Know more say less: Image captioning based on scene graphs,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.441150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.214435Z digest=sha256:fc358ff1d8ad8d7b8c0efbb39df3eaaf83d5758d96eb45f3c1362d2016268585

Observation 9192a4e7-23b4-4c75-aad3-0e2a41fc051a · outbound

This paper cites High-quality image cap- tioning with fine-grained and semantic-guided visual attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training High-quality image cap- tioning with fine-grained and semantic-guided visual attention,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:57.142395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.379950Z digest=sha256:1edbf316c0465a701f0e75190dc862edabcc49c0c72a185a6c2f0120f55fc74c

Observation 45cf7118-9da8-4cb3-b021-9ce38f332a1e · outbound

This paper cites Multimodal transformer with multi- view visual representation for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Multimodal transformer with multi- view visual representation for image captioning,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.954306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.450394Z digest=sha256:42f10d3a3b5845027fc97ea0ec148561be837c96eb6add31fa898e0a797445a6

Observation 9dbb5512-162e-4619-b482-2081258a674a · outbound

This paper cites Compact bidirectional transformer for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Compact bidirectional transformer for image captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.756704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.571278Z digest=sha256:92ce57b0bc53881cecb1f387e25a3ba52d3a3557cf9177b076653e7f98ed5885

Observation 58c688ad-e4fc-45fb-a40f-f41e3ad78348 · outbound

This paper cites Task-adaptive attention for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Task-adaptive attention for image captioning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.571741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.661175Z digest=sha256:71dc82454ce97fdaf5d59de7dbd761d2ae517753ffe01473fb032131d3d17e55

Observation b8faab56-4325-4121-a7fa-e0fc61873d7c · outbound

This paper cites Textual context-aware dense captioning with diverse words,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Textual context-aware dense captioning with diverse words,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:56.262369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:47.728248Z digest=sha256:23526c4c414284d2f5f081157ae61a9306479af7f68ccfd9d495ecad7ce288fa

Observation 46457da1-e0aa-4070-8373-e2adfac91838 · outbound

This paper cites Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.800623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.800623Z digest=sha256:090d5b336e8363d9177fb37c7f5474c96b53425e9cd72e56efbfa31bac1b5830

Observation a5722b42-588b-466a-be6c-900b74189f3e · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Rouge: A package for automatic evaluation of summaries,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.891704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.891704Z digest=sha256:03c64adea3a11a80c71e55718da80b3c69248f31116cfb3022dbfc769141820d

Observation b440dc69-0d02-42c1-a27d-3735867b32a8 · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spice: Semantic propositional image caption evaluation,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:47.956604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:47.956604Z digest=sha256:1481e2d3091b27ead4c7ba8ada87d5d7d5ba3eb3e38f8e6e4aedfe2863d09391

Observation 9266a018-e0eb-40f0-a5ec-d9b3eb341cf1 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.077301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.077301Z digest=sha256:69a83b41603f031d1b7380556998528721e5f0596b759f2cb9267dcb70acd02d

Observation ac99a651-3121-4d31-8668-a42d84296387 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Clipscore: A reference-free evaluation metric for image captioning,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.989883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:48.194780Z digest=sha256:ec3afaf37b8433949b26910323427d99a57c86eb88a3e6beb907c1e9ffae2b63

Observation 9e56f2de-0a44-46f0-8363-b3366ecad143 · outbound

This paper cites Tiger: Text-to-image grounding for image caption evalua- tion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tiger: Text-to-image grounding for image caption evalua- tion,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.753289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:48.326568Z digest=sha256:be07d48a4b98b69ea5bf144e4f68bcb44587cd91df9671f302920c185632e74b

Observation 473024b0-5264-48ed-91fa-011520487af3 · outbound

This paper cites Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Em- score: Evaluating video captioning via coarse-grained and fine-grained embedding matching,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.505044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:48.435085Z digest=sha256:2b12d55e6cfe7372ca7fcb15be0cc7dda77e7c5e4b98bf9526062a4d88e6c285

Observation 2bebff7b-075d-4dc3-9477-81a543bc9c20 · outbound

This paper cites Gpt-3: Its nature, scope, limits, and consequences,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Gpt-3: Its nature, scope, limits, and consequences,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.555233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.555233Z digest=sha256:34f31d64e37551390cc3d5dfa3b31fb6ff62dbdba45963db90da3944acfb6279

Observation 0e0d5cf6-ec6d-4216-9234-d04b0fb41152 · outbound

This paper cites Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bart: Denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehen- sion,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:55.274386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:48.636684Z digest=sha256:4515ead6d588c5a0b85481146fcc1bb42454ab1101b4e425f8e5f3fd6ff32d0b

Observation 1b91fbc6-9386-456e-9543-9953ab43f225 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Swin transformer: Hierarchical vision transformer using shifted windows,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.724702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.724702Z digest=sha256:9f8a6cfd26ec7cf513d25a06610e73b92a2d26f5fd77f497a2a9ad4f78477211

Observation 2f07502a-6d12-430e-8cbe-bf26a807362b · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:48.992310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:48.992310Z digest=sha256:c4a0e3bb57f794cfec7ff4eb666ff791285ce69104fcb76d7bb5750d4115b586

Observation e4c820e8-97fb-4d58-8dae-ad870c5dcfc1 · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.108576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.108576Z digest=sha256:4cbccf7ab732875c19afc12ce27e8ecc534d18df1614c16830d59ece7b33aeea

Observation 671a853a-d350-4185-a7bd-c000fcfee516 · outbound

This paper cites Slip: Self-supervision meets language-image pre-training,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Slip: Self-supervision meets language-image pre-training,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.996841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:49.294938Z digest=sha256:eb026930beb704b5006b511938e03a8610872330c4b45ad5f40344cb6c4d1474

Observation d5a17254-2f11-44e3-be44-f90bd97b8975 · outbound

This paper cites Pruning Filters for Efficient ConvNets.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Pruning Filters for Efficient ConvNets

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.439407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.439407Z digest=sha256:ce5f25ec6417d779fdcc806fff4b7cc1c50277aba6994a4e306d4c0539964cf9

Observation ede5327c-922a-4ed8-9555-cb7fb26c0b0c · outbound

This paper cites Vision Transformer Pruning.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vision Transformer Pruning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:49.586224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:49.586224Z digest=sha256:53f3c8b3d3dde7f7300f737af877bc224989dc199c5306c2ec7033249c91309a

Observation df4b184b-157f-4be5-a4cf-3c00dbcc0f2f · outbound

This paper cites Post-training quantization for vision transformer,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Post-training quantization for vision transformer,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.652256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:49.787638Z digest=sha256:8716d260deca16e4e6d633c7acb3f8eaca558806e01fd1cd3870579d08844c58

Observation 0505ab65-2ef6-496c-98ae-c4ec38866818 · outbound

This paper cites Tinyvit: Fast pretraining distillation for small vision transformers,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinyvit: Fast pretraining distillation for small vision transformers,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.444617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:49.905742Z digest=sha256:7aa6c4b9463ca9cdb133a9d3025ca9a682da1a9d330b6e8689ffc750e1dbff9f

Observation 0b0011c6-d04c-4753-abb2-3a5bdf78d27f · outbound

This paper cites Tinybert: Distilling bert for natural language understanding,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Tinybert: Distilling bert for natural language understanding,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.259692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:50.076329Z digest=sha256:38cf0290060c79b98c438876e6a26ba77c7e26b9212f69d148eef0680a208632

Observation 2b89c3d1-403f-4093-b9d5-8456c1de0eb0 · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.162524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.162524Z digest=sha256:f048bd448de81c08c595546da3e3da7406616895d8524544288f7e291ee5acde

Observation b39e71e5-76e7-4887-a9a7-cdab83f87c18 · outbound

This paper cites Training data-efficient image transformers & distillation through attention,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Training data-efficient image transformers & distillation through attention,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.303885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.303885Z digest=sha256:b8ad9fcb858ff356fa0688c0c954d5b4fa73f153fd0c04fddbb0000e4319b4ef

Observation 57e846a7-95fc-4f08-b549-f0006eaa6703 · outbound

This paper cites Attention is all you need,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Attention is all you need,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.460291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.460291Z digest=sha256:e35a197eaf1cd1d23757f2b7abc0cd2cf8b61c21c78e3e7e3b444ced1ec9f97d

Observation 507da7cf-40af-4f7e-b50a-a3421316f702 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:50.608287Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:50.608287Z digest=sha256:08bdb4af0c74bc5a2ea9d5004c6fa7a45006bc750046ac8249bdd6acb43cf9bc

Observation 6b05c5f1-1756-47bf-91a4-56717f2c1ae2 · outbound

This paper cites Bag of tricks for image classification with convolutional neural networks,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Bag of tricks for image classification with convolutional neural networks,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:54.104734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:50.761420Z digest=sha256:257d4f64498e093f2ca101c4ab0a33a9d84ae9163c6553b1bac1b74b30a16736

Observation 4af3bb92-6134-416d-b4aa-62c6937a5550 · outbound

This paper cites Making convolutional networks shift-invariant again,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Making convolutional networks shift-invariant again,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.800069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:50.886873Z digest=sha256:63df12da49cb083d3a49fc4c51adfc0c2a4af7977494f2ce254c201a394703c3

Observation 20d691ff-b298-4771-935b-70f04e205385 · outbound

This paper cites Discriminability objective for training descriptive captions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Discriminability objective for training descriptive captions,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.570868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:50.986021Z digest=sha256:4c8d3fc02e8d56aa9c263b28a4ab76311924f3ff2f0f4ce0d75ef0c98454a3c5

Observation 07f51bdc-c850-49b2-81e8-fb98b53f422f · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training ImageNet Large Scale Visual Recognition Challenge,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.056134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.056134Z digest=sha256:cb3c7ba13b6e5657c703a36208270cea2f091d04508365923eec1a3e16b518dc

Observation 2ba5db4d-d95c-41f0-a492-b0611a25abff · outbound

This paper cites Framing image description as a ranking task: Data, models and evaluation metrics,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Framing image description as a ranking task: Data, models and evaluation metrics,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.266631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.171017Z digest=sha256:ccb33903acf31d937bcf5457e15aed6637654c7d26d7fef18aebb292f2dc22a8

Observation fa8cb4f5-4a9b-4e46-b2d8-af3aa3151a8b · outbound

This paper cites From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training From Images to Sentences through Scene Description Graphs using Commonsense Reasoning and Knowledge

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.241514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.241514Z digest=sha256:f7dabd63d2f35aab0182257369ca047d114be7a3e5b57fd3c04f2ab1aa5edf2b

Observation b5ca7f10-23d0-4445-9220-275089ddd46e · outbound

This paper cites Consensus-based image description evaluation,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Consensus-based image description evaluation,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:53.042948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.305407Z digest=sha256:ac01e1faa8ab9925a43d25e19f80aaddafebbf62c4e535811bf1e31b8d9a110f

Observation 0433243b-fd85-4fcc-af24-11e652ff5144 · outbound

This paper cites Foil it! find one mismatch between image and language caption,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Foil it! find one mismatch between image and language caption,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.700175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.395791Z digest=sha256:96cd103e0d0d55f723ab72e09977ce02d6a4e2437f11fcc8476ff27f53cc3ecf

Observation 6705d649-e7d7-4e3f-9634-a4129066491c · outbound

This paper cites Deep visual-semantic alignments for gen- erating image descriptions,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Deep visual-semantic alignments for gen- erating image descriptions,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.538573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.477940Z digest=sha256:feb0e6708077b805aa353c02a7fe6c6887e48c971a6eab56a170fa00261dac1f

Observation 24f8f060-4a72-4d69-a4fb-4b7d196df7b4 · outbound

This paper cites Vinvl: Revisiting visual representations in vision-language models,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Vinvl: Revisiting visual representations in vision-language models,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.392262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.554693Z digest=sha256:a8f269850c7490291b35cd250e76cde776d5ffde01a08aa09285132777a3afe1

Observation 8dd056fb-e566-42b9-a6c3-2a426bbf0ca9 · outbound

This paper cites Adam: A method for stochastic optimization,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Adam: A method for stochastic optimization,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.258827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.640789Z digest=sha256:b9269839b0762bac8cce6a2125d162658a51e17f0a283cac5392fcfa98374a0a

Observation fc4aa0b8-2168-462a-8dd9-58111a0fc563 · outbound

This paper cites Improving image captioning evaluation by considering inter references variance,.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Improving image captioning evaluation by considering inter references variance,

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:18:52.129394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:18:51.705906Z digest=sha256:6d4b3d41eb60ca51554567b1e463b43fd98ff949e5b2097d822ea3a42376c10c

Observation 380b0f08-6362-4ce7-8203-5776db952ae8 · outbound

This paper cites Concrete Problems in AI Safety.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Concrete Problems in AI Safety

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.809319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.809319Z digest=sha256:0d3e1e0e9be87eb8e5af107638b68051ab5adfbbfc597ad2e9daefeca0b4b2b8

Observation 61a01979-bf04-494a-9ff6-eac046908ce5 · outbound

This paper cites Spontaneous Reward Hacking in Iterative Self-Refinement.

L-CLIPScore: a Lightweight Embedding-based Captioning Metric for Evaluating and Training Spontaneous Reward Hacking in Iterative Self-Refinement

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T18:18:51.876640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:18:51.876640Z digest=sha256:64ab8584cdcafee58f7d648920ffeedb8a67e989beeb9b534f4ce4279e6b4745

Pith citing papers

No inbound Pith citation observations are available.