Pith. sign in

Paper Citation Record · LEDGER

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

As of 11 August 2026, this Paper Citation Record lists 63 of 63 outbound references and 2 inbound Pith citation observations for arXiv:2508.06125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.06125 v1

Coverage vector

measured 63 of 63 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:40.901060Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T13:13:57.599970Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T05:47:41.745200Z

Reference resolution

63 of 63 outbound references displayed

  • verified exact2
  • verified fuzzy27
  • unresolved33
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5e3188e3-386c-4c21-9782-4ab79c49b35a · outbound

This paper cites GPT-4 Technical Report.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.378135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.378135Z digest=sha256:cf72f64d0d222e2b1a774024069395f2f849fdff8598dd47355ff325624495a7

Observation 9f187492-904b-4d1b-ae95-48249658e748 · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Meteor: An automatic metric for mt evaluation with improved correlation with hu- man judgments

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:48.647640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:33.457218Z digest=sha256:06e3ceb2a6851cf213a2173e24df4cdbcf9532f4a91a2a97259bc94a9ebffb59

Observation 81029bf8-c384-4591-b8db-6149b0d2cb01 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.568450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.568450Z digest=sha256:cf254ac93a04acacfc00ccb3745ea47286d388b67975091f38ae74c9298611fa

Observation da60b3e5-d993-46ec-b9e0-2c35daee4b04 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.657104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.657104Z digest=sha256:d1844b483e8d34b23e8df2326748a1c1e877fe88e85ab52f73527dc17a6387dd

Observation 54adc342-0533-4de7-98ec-c65214ca9fa3 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.759054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.759054Z digest=sha256:fe5243f5c944e3986325cffa3ad7f3a87f9043857d438908af6efa4d0e0110ef

Observation 91375458-db5f-4dd4-b3d5-57eaf63c1781 · outbound

This paper cites Meshed-memory transformer for image cap- tioning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Meshed-memory transformer for image cap- tioning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:33.847904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:33.847904Z digest=sha256:d05cac20b59960abb67335a5195ccc8e8feab764a55b759a47831aaa27abff58

Observation 7bd57807-a7c1-4248-af0a-f4f95a150ad8 · outbound

This paper cites Cross-domain image captioning with dis- criminative finetuning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Cross-domain image captioning with dis- criminative finetuning

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:48.411911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:33.971679Z digest=sha256:214ba2ca7dabaa02d865126aaa7f8e8d99dbe0b630e929763b1878fa575a45a5

Observation f4df095d-74b4-4c46-ba70-32982931c397 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Benchmarking and Improving Detail Image Caption

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.074648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.074648Z digest=sha256:9554a337a70a601356c34a9dae5816ead67a90695119dea07c6e0b66746bc09f

Observation 20e663f2-2d1a-4896-be48-dd4eba14e8d2 · outbound

This paper cites VILA$^2$: VILA Augmented VILA.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning VILA$^2$: VILA Augmented VILA

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.166423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.166423Z digest=sha256:e4e6c7ffca3b8005bf605543f8a3e041c585fbc23a357509ff4d1714f8a54d95

Observation 3bb6d37e-5fe3-4eee-ba19-fa80475479ff · outbound

This paper cites The Capacity for Moral Self-Correction in Large Language Models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning The Capacity for Moral Self-Correction in Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.281893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.281893Z digest=sha256:5e13e1fac60d6f135e04ccbf9881f2bfee41f715fe0388ffe197f2ea217ab493

Observation 63882c2d-434b-4b7d-ae1f-80fc475e7235 · outbound

This paper cites GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning GLoRe: When, Where, and How to Improve LLM Reasoning via Global and Local Refinements

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.400113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.400113Z digest=sha256:0c8d9e3d5335ed73312b477fdec53827b0a48a59d8c8bd80dcabfbbd69398a6a

Observation 2540fa4c-3e07-4c8f-815c-a1408909e327 · outbound

This paper cites Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Self-Correction is More than Refinement: A Learning Framework for Visual and Language Reasoning Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.507296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.507296Z digest=sha256:558acc8ddea1284b25e3250fee768f01cca43d212e31851858b1f498eb4b2a6b

Observation 06beefce-f4e2-438e-98c5-055501d80fba · outbound

This paper cites A topic-level self-correctional ap- proach to mitigate hallucinations in mllms.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning A topic-level self-correctional ap- proach to mitigate hallucinations in mllms

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.625129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.625129Z digest=sha256:69667acf4e1095670709dbf52311de7eabffeae82b9f4aae89d0a83081c5ade8

Observation b91d7e3c-1ff0-4f04-993b-0eb487e04923 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.731038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.731038Z digest=sha256:cae1e834163525ef1cbbb8d5b0050088731bec3466cd1e6167586ce383e1d99a

Observation d374d9dd-bcc8-483e-b97a-32676950de1c · outbound

This paper cites Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:34.844196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:34.844196Z digest=sha256:866996f8ffa9b7c83ccd4975aee7af7660312ebfcfc53780fc02da8daccc54ed

Observation 37921116-0a10-4ca5-ab78-f61a39f1545b · outbound

This paper cites Lora: Low-rank adaptation of large language models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Lora: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.005045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.005045Z digest=sha256:766ea1884d6641eb8f644c1388229eeefcfe652e90bf430fff72d9aeeb6af55e

Observation 00203f23-2d0f-4019-9c19-e14d37591905 · outbound

This paper cites Large Language Models Cannot Self-Correct Reasoning Yet.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Large Language Models Cannot Self-Correct Reasoning Yet

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.116852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.116852Z digest=sha256:e7c5b001ce71936a33282446bd81eda969c369ed849b5b42f25a3ebd7b2861ab

Observation 3a743dbd-9b49-4019-a201-28a5b9079e57 · outbound

This paper cites Attention on attention for image captioning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Attention on attention for image captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:48.204818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:35.236922Z digest=sha256:82d8530e06d1ea201eba990134d4cfbf02c650b31da17c72eb92930ea243d28e

Observation ea6d09f7-4d09-4c89-865f-0c382b07344a · outbound

This paper cites When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.361067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.361067Z digest=sha256:0562ddef13e47a643684ea1436287d5b9ec76b34f79073bdc51c923516daf75d

Observation 4685203a-1490-4f17-bcf3-cebe0123f324 · outbound

This paper cites Training Language Models to Self-Correct via Reinforcement Learning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Training Language Models to Self-Correct via Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.455107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.455107Z digest=sha256:0374746c7d9098e6bd9be06479a8627c6c9326fd1e8e6b2fc017803097b85345

Observation c4a0e725-8dca-46cc-a87e-50b4e00b0c9c · outbound

This paper cites QACE: Asking Questions to Evaluate an Image Caption.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning QACE: Asking Questions to Evaluate an Image Caption

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-08-05T22:58:41.513395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:35.624153Z digest=sha256:5213d93f4a5cb5338aeae34e645551fffd715a2333d572be8a195844de840ef5

Observation 799cc422-8eb3-49af-a754-0a55aa35e79b · outbound

This paper cites Fleur: An explainable reference-free evaluation metric for image cap- tioning using a large multimodal model.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Fleur: An explainable reference-free evaluation metric for image cap- tioning using a large multimodal model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:48.016979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:35.769242Z digest=sha256:35bc4c5c587d9fdfb3e4c4218acf532ddad2d9ee10afb58d33202fbd678709c2

Observation 0a8e2c5b-f20b-4ba8-bc43-6cd5433ca931 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:35.859637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:35.859637Z digest=sha256:169b22a14f81f561f81ca753c21103a715c8e8f7ee34b0394c5045ca9212ca28

Observation 4744dc6d-d104-4962-9f6c-e6cf2ee0aab3 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:47.858272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:35.970216Z digest=sha256:7f5810ea2a7571d06496754995be32fa9461ba41b2c6d11a174d5020edc1282c

Observation 243f5353-fe70-4540-9a9e-2fe480bc6f85 · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.145548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.145548Z digest=sha256:0f9bce93e6ca2d0f81c6e6d0b3a019d2c9c98213f29640caa227fb75abf69763

Observation 75eef57d-30b7-43ab-9547-73adf55f69ef · outbound

This paper cites Mon- key: Image resolution and text label are important things for large multi-modal models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Mon- key: Image resolution and text label are important things for large multi-modal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:47.661173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:36.281087Z digest=sha256:f7ae58f3557438e66ff2db7bc9e1d0c8261f334370dbc916166a7cefa06eda9f

Observation 2799051a-0be3-4f43-b590-fc8bc919b5c3 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.402785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.402785Z digest=sha256:f10963292fbe46a4e467b90b3ab5825f3acef65e0a3eeb484012ccdefe1bc3b0

Observation d843b8d9-6d8b-4100-b8db-cb4b2d9a44a8 · outbound

This paper cites Attention correctness in neural image captioning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Attention correctness in neural image captioning

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:47.435775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:36.515020Z digest=sha256:71faba13494298c7532db4a31af4982092822ba2e1bd9da20868aaa2a2e28c0c

Observation 505f2ecf-d03b-4ad9-9120-e598a5047645 · outbound

This paper cites Improved baselines with visual instruction tuning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Improved baselines with visual instruction tuning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:47.252670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:36.641783Z digest=sha256:0e2dd451372d6441477e951e0cd899427f9bcf9b63b4e45f26c5758fc0709734

Observation 37e799a1-e9e0-4b04-881d-a42c76960982 · outbound

This paper cites Improved image captioning via policy gra- dient optimization of spider.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Improved image captioning via policy gra- dient optimization of spider

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:47.053072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:36.770853Z digest=sha256:689ba03316d64a9f2993064b2a587b903405da3c63992c903ded9b9c91621667

Observation e91735c6-df67-4819-b970-c7a0857b77f5 · outbound

This paper cites OmniCaptioner: One Captioner to Rule Them All.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning OmniCaptioner: One Captioner to Rule Them All

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.852405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.852405Z digest=sha256:544dbea1c1eda51d328cbae858d8575e16fdd58632f181a917e5c63d0379f2bb

Observation 96c8ce27-17b8-40b4-ad4a-746c5b2b2ca1 · outbound

This paper cites Self-refine: It- erative refinement with self-feedback.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Self-refine: It- erative refinement with self-feedback

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:46.771264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:36.969470Z digest=sha256:675a7ae24bbeb388be0c1f5482a6c64dd53dc04a927d3357072f9d26e1ea1c88

Observation 26230f63-8435-458b-aceb-cb054c5dfbb1 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning DOCCI: Descriptions of Connected and Contrasting Images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.077028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.077028Z digest=sha256:4ab42a06cdcb2d29b90f36e24c2d3bf5312cb79d2edfbd220eaeb024920e6924

Observation 99e96986-63ee-4bb0-836c-0245245c421e · outbound

This paper cites Automat- ically correcting large language models: Surveying the land- scape of diverse automated correction strategies.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Automat- ically correcting large language models: Surveying the land- scape of diverse automated correction strategies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:46.525560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:37.169905Z digest=sha256:fc4328d34a676a117946ae464bf287b267a3b1aaee4d12e998ac0bf05cf57401

Observation 75aa1382-222b-4f8a-9222-f898f73c193a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Bleu: a method for automatic evaluation of machine translation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.288984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.288984Z digest=sha256:63822f74335eab7e5c42605f08fa69db37b3539d465d93dea1c5ab0a93e95d4d

Observation 3041647a-bb19-40b0-b259-83ba4ea9c0a0 · outbound

This paper cites Connecting vision and lan- guage with localized narratives.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Connecting vision and lan- guage with localized narratives

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:46.223990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:37.415476Z digest=sha256:2bf7fc9a5ccc203b3bf460ff190edfca530f4055deacbdc4d4c63dd0b2cf1182

Observation 91ebfb98-f1e2-4576-9d57-331776b8c3f2 · outbound

This paper cites Is moral self-correction an innate capability of large language models? a mechanistic analysis to self-correction.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Is moral self-correction an innate capability of large language models? a mechanistic analysis to self-correction

Reference 37

Resolution
verified exact
raw_fallback, observed 2026-08-05T22:58:41.231786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:37.565099Z digest=sha256:91f37c83e5527b7ef724880c9e4bf200057ae995c59d4919bfb6d6ff8a7ca338

Observation 690e37c7-23ff-4ae4-8dff-7e1464162d0c · outbound

This paper cites Recursive Introspection: Teaching Language Model Agents How to Self-Improve.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Recursive Introspection: Teaching Language Model Agents How to Self-Improve

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.669243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.669243Z digest=sha256:c73a4630aa626aa629772721df169ca9a808304e18ec7d3f81bcd84310351e40

Observation c32ae3a1-84b9-4e9a-b43e-c5d0b3017141 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Learning transferable visual models from natural language supervi- sion

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:37.816834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:37.816834Z digest=sha256:d7a998a08e78c4f4bbace0ec5b36f5bd595e0466a153b59dc5b2528db067aaea

Observation 9dfca791-a63c-4c28-9642-96780ec9c911 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:45.984325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:37.939922Z digest=sha256:900ff625da807efe7fb361ac22c402f742c365c76363c2d55488053564e0c036

Observation 08c782bd-6a05-4080-a722-53ccfdf62425 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Direct preference optimization: Your language model is secretly a reward model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:45.681047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.045115Z digest=sha256:9135c5a296237e3a4738f509057f8646d2f09a6adbbf92723322cb38c3640cca

Observation c9852579-f688-40ca-bf6f-6fd376369061 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Glamm: Pixel grounding large multimodal model

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:45.453769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.171749Z digest=sha256:5eb429f2705137f601fa0d2b982b4f52a1ce6dc76aec2665e5d44471e5771dbe

Observation f95dadf4-f361-455e-9a76-a5ed4739755c · outbound

This paper cites Positive-augmented contrastive learning for image and video captioning evaluation.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Positive-augmented contrastive learning for image and video captioning evaluation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:45.194254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.323923Z digest=sha256:1a6be5ad256c4cd32896e5ebd4e4b55b3ae0c692d25dbe624f1cb8ce3b5df793

Observation 33b2b492-5123-4879-8039-604b1a969401 · outbound

This paper cites Bridge: Bridging gaps in image captioning evalua- tion with stronger visual cues.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Bridge: Bridging gaps in image captioning evalua- tion with stronger visual cues

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:44.930966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.472255Z digest=sha256:a99e201422ce03dc79c866f7bf4c4de9ef643c1ac587d613b43138899638337a

Observation f87fb521-c4d0-4be9-b65c-03c0c6660ccd · outbound

This paper cites InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning InternAgent: When Agent Becomes the Scientist -- Building Closed-Loop System from Hypothesis to Verification

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:38.568367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:38.568367Z digest=sha256:8aebf213ac3f23b0db073f37f9e3ed488e1aa97c46917758133c42cd504c42bd

Observation 681ebff4-3477-47b3-a968-b5b0caffed6b · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:44.663601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.680862Z digest=sha256:87002646db0162cdf6a0bca5d6d5c06a26689eb51e0babd342606be17fdd554a

Observation 21eca74e-7c51-4579-b5c0-ab70b2cc16c6 · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Cider: Consensus-based image description evalua- tion

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:44.398320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.823823Z digest=sha256:b092bc4365e0dc4fdcc8452d6840f93ef26d728d3a2eba8c68baf6d812b01403

Observation 6ee465db-8c59-4529-82d2-34a60ff074f9 · outbound

This paper cites Show and tell: A neural image caption gen- erator.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Show and tell: A neural image caption gen- erator

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:44.133064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:38.913563Z digest=sha256:cb2bcd4f89c387582bec570494e128a80226dcc8cc836b480c6ff3e699a462f8

Observation c70be9ec-30c4-4eca-90c2-f4446829c581 · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.078760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.078760Z digest=sha256:ce1b8f6ed3c9428a094535d8f2bc8bb6d24511ab0d20830477d3881529bc19c6

Observation cb6864a9-ffa5-41ed-82a0-34c08ddeb147 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.203557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.203557Z digest=sha256:ddacc3244c0dd00270444931687617f1feb848cb5e6d3635ff68a8a64c2a330d

Observation 50de6f54-467d-4a06-8f6c-716158020434 · outbound

This paper cites Caption Anything: Interactive Image Description with Diverse Multimodal Controls.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Caption Anything: Interactive Image Description with Diverse Multimodal Controls

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.332921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.332921Z digest=sha256:49af09cc1e4508ed06a545096ec109b674c0de516a91a829dfb1593986c25645

Observation 2a05045c-dd8d-48e8-8b85-dbd2e8d22b01 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning CogVLM: Visual Expert for Pretrained Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.456853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.456853Z digest=sha256:0e864b1d4195e0de439db2790b9787fc0f3e48d2c7bdf294f5fb704fc2f5febe

Observation a0a86057-9ea6-4940-a317-98ab80e2f66f · outbound

This paper cites Generating Sequences by Learning to Self-Correct.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Generating Sequences by Learning to Self-Correct

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.580049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.580049Z digest=sha256:403cdfbf3c83706bf6ca35118e74a2ef4a1f1378e1b5f6c0b1cbc56c789fbb1a

Observation 5ce173ae-a26d-4bbe-821d-8178f49ad673 · outbound

This paper cites Painting with words: Elevating detailed image cap- tioning with benchmark and alignment learning.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Painting with words: Elevating detailed image cap- tioning with benchmark and alignment learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:43.799450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:39.717372Z digest=sha256:21e6f966e6d39d6a19e38190a16b9f3cb8791e007ec832cd08a02698018ff93e

Observation 06973013-4ceb-4e65-9981-10885271bcc2 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning From image descriptions to visual denotations: New similarity metrics for semantic inference over event descrip- tions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:43.538357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:39.861379Z digest=sha256:c0194aa73788c8937bb650e2111029bb2e327344309851d5de5d2a3297653b86

Observation 286c0451-cefe-417d-b5b2-539ff841a4cc · outbound

This paper cites Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Dolphin: Moving Towards Closed-loop Auto-research through Thinking, Practice, and Feedback

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:39.953810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:39.953810Z digest=sha256:5718de333568cb558295ad2322950293d4ac3de58f2f5355d09fc8920d357013

Observation a4143063-826d-4d87-8f4b-d3928d2c9a17 · outbound

This paper cites Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Llamafac- tory: Unified efficient fine-tuning of 100+ language mod- els

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:43.308003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.076459Z digest=sha256:d250cfe80081d3f24e9f141efc6afd34bb964f2580b39c46f4f2666df17de189

Observation f019cea8-0459-4f92-9edc-ef2aa042c1d1 · outbound

This paper cites For relation evaluation, we prompt open-source language models to answer the given questions based on the candi- date captions.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning For relation evaluation, we prompt open-source language models to answer the given questions based on the candi- date captions

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:43.061158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.202037Z digest=sha256:c5da413cb65879187a761d4f0e20099856b6a25dbca60911946b06eedffe2143

Observation 428c3552-a994-4ef9-81e8-7a21593a6f4d · outbound

This paper cites We randomly select 100 images in DOCCI500 and ask 4 hu- man annotators to sort the captions provided by 4 differ- ent models, while considering both precision and recall.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning We randomly select 100 images in DOCCI500 and ask 4 hu- man annotators to sort the captions provided by 4 differ- ent models, while considering both precision and recall

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:42.872936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.354135Z digest=sha256:37fee4ded94dfd7e0bc1303744f9aeadb464982cf0ceb0b93eff8cddfbedb2d2

Observation c2004cf0-d76e-4d56-98cf-8b8a7def2239 · outbound

This paper cites an unresolved cited work.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Unresolved cited work

Reference 60

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T22:58:42.675978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.453530Z digest=sha256:2309277d366697680496a311349dcdbce4bd376750e18756791b071e3152d3ba

Observation fda26ee6-9fde-4241-a578-41c2aa1e3087 · outbound

This paper cites Same-Domain.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Same-Domain

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:42.425937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.599848Z digest=sha256:7033228554d1973c921aa4a099b318330c7652bc7cae997806403400ebf9748b

Observation 735b9408-4ed1-4bcf-a845-1b810bb084be · outbound

This paper cites an unresolved cited work.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:58:42.154224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.739659Z digest=sha256:3a6f454e5cf5c3543dfceae3f8e077b1aa7d11b4f13087701460e1e3f3752291

Observation 66ef2bd3-43ed-49b6-8f6b-afe224ba9a46 · outbound

This paper cites Because the training process includ- ing generating annotations for two rounds, the training time is relatively long.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning Because the training process includ- ing generating annotations for two rounds, the training time is relatively long

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:58:41.900996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T22:58:40.901060Z digest=sha256:2f89411efeb9b3f3168a0e8a591ff1cf367e03d143e78f9b8282382e553869e1

Pith citing papers

Observation 13681d6c-ab6c-4375-8891-fa98f517130c · inbound

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning cites this paper.

VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:23:28.464514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T13:13:57.599970Z digest=sha256:73c59dbd099110b01798b9812007af4d525d0f6933c954795b3cbcb26da3b51e

Observation 32df02c4-e8c5-432c-9613-d439311c6804 · inbound

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models cites this paper.

Modeling Complex Behaviors: Multi-Personality Composition and Dynamic Switching in Vision-Language Models SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-07-03T05:47:41.747057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T13:04:14.886733Z digest=sha256:5b532c26d8a006396a1d99be7b552c71005c87ca05b5da0bf3c6c389e9f11962