Pith. sign in

Paper Citation Record · LEDGER

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction

As of 8 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.22613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22613 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:08.869281Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90463c61-bf5e-443d-8851-0b8611e30f46 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.224914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.224914Z digest=sha256:923f2d51d505d177017828121fa97c6d4b7a143f113d98a4168b67be73feb895

Observation dc5ff3fa-1c18-47dd-b333-692732d69e67 · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction SPICE: Semantic Propositional Image Caption Evaluation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.296549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.296549Z digest=sha256:6250946b0f2e94f233cf8ee9aa92f4c8d7c1e67de5e69c0ca03f6a3df75a2e31

Observation 714fbf1d-f07e-4025-aea3-127cae2dfbc4 · outbound

This paper cites Qwen Technical Report.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.411249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.411249Z digest=sha256:707aee22f2f225b70b9328a87ec7a836f208f5cccb6c3faf5feb89504783db1c

Observation dc6bcb00-ea6c-45ef-a8e1-9e2c5364cc79 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.514104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.514104Z digest=sha256:48926f50eedc07e08bfae51b426aef592a9ddc344ec426c2dff229802fde29a7

Observation e803da03-b8c5-45fa-9602-7b9c93be180c · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Hallucination of Multimodal Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.609812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.609812Z digest=sha256:6ae1d7a35d9d6eb860af77b6a360c24b6a32b8fd6bbb58e0f468fe0ac9017af6

Observation e97fa7e0-4965-4cb8-8b22-5fc9a85af19b · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:08:10.300178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:08:02.710978Z digest=sha256:938349b0c7fe3f0276cf98fd2d1eb052f5448cd905b3f80ec241e4aca3a01e78

Observation e059b8d5-4d35-4768-99d0-84a0df065497 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.797253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.797253Z digest=sha256:60caaf00247dfecf1e79ff5bb764897408a3b3add9b3a163dfb2e22aa254c8da

Observation 5cb60e30-a5c0-4049-9cb6-7d47de177c7f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.914321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.914321Z digest=sha256:7cb9b6958823e8538ef704a4bfc02ff6f979c049fc571b4608e9394c8d8225f9

Observation ec6b558f-1628-4982-9668-2ca3c7f3a36b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.022462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.022462Z digest=sha256:f7698630f62fd81a5a513046191ec000f353e388e59d7288cf63f3af3d545e04

Observation 285fe182-3d78-4df4-860c-79512ddbfdd4 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.153714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.153714Z digest=sha256:96de7c9a8420b1ed5de7aedf0805778050efe50a2f5d9112dd7f863e72062c18

Observation 93f5842e-428b-43d3-8fb6-f97863357a38 · outbound

This paper cites A Survey on In-context Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction A Survey on In-context Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.233342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.233342Z digest=sha256:25b4866a00cad761ae1f463b8f76bd84927de51891b59ba127e5fcf05056cf10

Observation a0697e11-d0e0-4a02-872b-e8111bb91916 · outbound

This paper cites Improving CLIP Training with Language Rewrites.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Improving CLIP Training with Language Rewrites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.303022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.303022Z digest=sha256:0715174f0b8942b9b6e64a80cecd89efed5a2f7e58fecb2950d0e77ac4953a90

Observation e1c9293f-1173-4e62-b229-6156df03b9d9 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.374847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.374847Z digest=sha256:3e7fcbf3c40299da6738c898e170d2e160f1cf773278086bef191b9fa9b384b2

Observation e9b1095e-54f8-46ee-9570-65993b2f0c2d · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.447542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.447542Z digest=sha256:50a23bfc471be0669570e49a568cf0723247c09851463b1d7d8d04b8cf4d93ec

Observation f17a7886-2c4c-4fc3-b881-22885e2569bc · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.524335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.524335Z digest=sha256:8b3d641b9c69818e067b82ec1a16db9fbbffd86a71948403ac1e1ae8dc9c788e

Observation a683d7db-d212-4e2f-92ee-460c5bb7e6bc · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.626178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.626178Z digest=sha256:1cf5f498efb8982f044fcd96f5a1531cf87bb344da99293115d5b9fe70e86fb2

Observation 8e721d09-db49-4d03-ad35-f21644d1adc5 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.741011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.741011Z digest=sha256:1d354f95a6ab3753fe5bc07f9ca2fd3b1884cbc46eec72fa726e823c2ca286c1

Observation 67873c65-c2b7-47b4-93ef-2441ad0781a2 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.841863Z digest=sha256:bdc0f86d168826b5a5ffd522b49048b435357f711e802ad6e848a5123d819085

Observation ab514049-398b-44d1-9f7b-4630afa36023 · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.913503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.913503Z digest=sha256:67cd8be2270294094425d49aa75a53d0470ca8b98154ac0d05c6e30ae3c047ff

Observation 65f33dcf-9ea4-497b-a508-52d258668a74 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.998066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.998066Z digest=sha256:4528b2980ebb9ff000ca1327d0543c0731a75b0e0fec12a228501fb98ccaff94

Observation 9e7fa958-1490-4002-81da-fac08bbf48bf · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.091130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.091130Z digest=sha256:26545327efea1587cab1efab431240a4a089aafa29bd82f9d2f871d0c7bca3e8

Observation ae4bd084-619c-4959-a82d-421bd37155db · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.205365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.205365Z digest=sha256:ddb5b1a9212d57701e4759efe6e240666388f1b3d94c5623e63872019306380a

Observation 4af59865-4ae8-42ba-a9c6-40fe3663865b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.303127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.303127Z digest=sha256:2778c4a247727e8b502efd1cdc192e3fceea4745da627558c09f6979cf3cc476

Observation a14b41b1-eede-4c43-b21c-bf8d2717c0e5 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.386112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.386112Z digest=sha256:946c2fe3dbd8cfb16457bb90ba9e8196d8e4c96364e3c4829b58ea8495ccbd40

Observation 50aa06cf-5d1b-49a6-8515-95435e9fd6d2 · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.479059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.479059Z digest=sha256:f0bedec61e4d1ddbcb994400c3177372e19fd237b76239026871b93a8b37e325

Observation f82f29cd-725b-4492-89dd-b8a0c02d954b · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.565373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.565373Z digest=sha256:783cf6cb66e3b4f3730a3223785b258ce702546dadc636703d83768c39391698

Observation f79b9c52-9e3f-46e8-9e9a-96587df0c713 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.662605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.662605Z digest=sha256:40b045fb06a4d1c5a1d72dc5978133e5df69283596255aff672a1f86b3b8489a

Observation 25002f22-4f2d-4fbf-9222-f49370bb2c88 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.804019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.804019Z digest=sha256:cdf368eea27a8ec91eb94fd0484cb7c74d91d240cb667a1f7459a5fc2c25a34b

Observation 27ef73e7-64f7-4ca8-92d9-a7a406674719 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Improved Baselines with Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.893325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.893325Z digest=sha256:6dce29d021acacbe3a8a27649bdbce37649d1040ff66e90e8af0158ebfe8dd56

Observation f3c8e907-8dcc-4ff2-93ed-c5653997fe54 · outbound

This paper cites Visual Instruction Tuning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.034601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.034601Z digest=sha256:57cacb4b4d47042404ceef14773cc9794f620b0497705aa2f84f707ee65c2902

Observation 35c61650-6da7-4968-a50d-1fe8be2d38ea · outbound

This paper cites Decoupled Weight Decay Regularization.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.189909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.189909Z digest=sha256:8749d1b024f098fa1ffea176d3ea70e79f2cf7d9746036c5310c32cadd823479

Observation 050d1fec-5a17-4cbb-a3dd-6661b90d05d4 · outbound

This paper cites Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.395780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.395780Z digest=sha256:12872458e13b5f8424d36777bd4b0bc34662faf60a36ff67ab78f09d0d143fbd

Observation 7fbd89d8-0707-44bc-b076-3e5942c8d3fb · outbound

This paper cites GPT-4o System Card.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction GPT-4o System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.567642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.567642Z digest=sha256:490fd111461c58f5b650d2480c3c890ecfcbe17fe2a91f00aa7e1bc760fe9641

Observation e12b594c-8055-4723-88d4-332c0e529590 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.741641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.741641Z digest=sha256:fba6e154785538dca7c774be4aa56e090f741d466e2952feb1d4513a640c9473

Observation a1f97d57-0233-458c-8d8c-b6f5ad7bcb87 · outbound

This paper cites Training language models to follow instructions with human feedback.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Training language models to follow instructions with human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.893497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.893497Z digest=sha256:61edf805b1e00c5c7ca2e6dee27c9c756a1aa6e5a094ce6070e1b189cf4dbbfa

Observation 9e3b42fd-d060-407b-961f-0a91638baa29 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.015674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.015674Z digest=sha256:2edb767796f01bdc173ba3d15919a3063522182d3c3d61f2fc08f1bec776ec81

Observation f8243aab-30e4-4d78-8449-60e07b5b33d0 · outbound

This paper cites Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.657590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:08:06.136121Z digest=sha256:05206e81843de72dd776e88f6a91c6f7cb56c25ced7afe5b78d7054aa29b91a9

Observation b96e1fe6-ade3-4e55-99f6-14941409d313 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.252023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.252023Z digest=sha256:55fc4915ed16f48beacc8c8d390d9a69a0c803619b88300a50bb6d67be7d9ecf

Observation db01e3cc-8589-40c3-811e-42278e299bee · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.370126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.370126Z digest=sha256:33c4556b1dfe72d9b5c7347e2ab3b32524b14d06afb61edda19154977de26336

Observation 0bbf7527-2e0a-43dd-8862-846ec2812208 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.501876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.501876Z digest=sha256:f53e437ac9912274c124e6689eff460a4d6db3fbaa42615f370633ed61067973

Observation b91d8922-fafe-4c54-8071-bc1d40044a68 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.641934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.641934Z digest=sha256:61daa2783aadb880e6871c4a42f580b3ae4125f5d9402889288b08dd4ac7caf8

Observation a85e9afc-a2df-4d8a-9b1a-719b8539f95b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.759252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.759252Z digest=sha256:f9173131e05c623733aa4482acfad6e9981d99f465323022360a9712ca604156

Observation 24b18474-4d6c-4d2b-a31f-3b23fe409bb8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.905429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.905429Z digest=sha256:8063ff9997aa178ab783160aeb6f20d3756d8f1f72bafb7e1ce6e7b50944c5ed

Observation 1fe5ad3a-fbac-417a-9cd5-dfb1c964319b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.021522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.021522Z digest=sha256:9aa4aca5020ccce2b3f90a3c58d8620729d9e3d2f1971f84cc01750ac48836a7

Observation 1f58cc58-b22b-45a3-ad0c-4314b2fb9870 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CIDEr: Consensus-based Image Description Evaluation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.125212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.125212Z digest=sha256:dd33e9a9af88f799c0cc804627ad50fb4f1675210448a755bda1eea51a2d9e02

Observation 30f5fc9a-3a27-4395-af93-e472c428b700 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.239579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.239579Z digest=sha256:2b9b3659be9fe09bbaa08d86090f3dfb249c5e5192c15109b3b08316eae38b9d

Observation da252b43-a893-4478-abe5-2eb5a55d4d6e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.365190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.365190Z digest=sha256:f313b27006b531d6cc61ccab072977a686d19f36538f658d31d5cbdf0f9e9b38

Observation 10e2b154-614f-429e-bd7e-aceb10d84581 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM: Visual Expert for Pretrained Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.562423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.562423Z digest=sha256:26b82a76942998f9b2ad065086636f25abf80be5219e8bdf758e10affcc79064

Observation 0946a2de-47c2-4566-92ca-7f82c54c7637 · outbound

This paper cites LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.404505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:08:07.687577Z digest=sha256:3edc36f998d33a932c50e06821bd07922c9e3f3b496443364300effcd7140aad

Observation c8d4c1e2-bb6e-428d-984e-240281321dc1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.783271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.783271Z digest=sha256:fa41bb444bb6f6831a8934d9ef732281e74959a7b83040ed2cffd6d569363068

Observation b48f0b6e-904e-436b-ace1-e253840de87c · outbound

This paper cites Altogether: Image Captioning via Re-aligning Alt-text.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Altogether: Image Captioning via Re-aligning Alt-text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.915717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.915717Z digest=sha256:e73ad3e4770713b46b3db594bc5926d72bdb81afc55ef57e6feb6242e2779895

Observation 4a1361b2-4b68-4d2e-adf4-9a1239c4f5a1 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.086271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.086271Z digest=sha256:b578784ff0620a132b4cae77dabbefd2bbd0afe24d561264780b9872e63d70b0

Observation e07b8e6d-61fa-4e78-9970-982a6a4e66c1 · outbound

This paper cites CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.129159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-07T13:08:08.229404Z digest=sha256:bad8d6a936ffb6a91b19430e0d987a5ab146af10f668de7a84d2d45812d1c2b9

Observation 264c1c5c-0e8a-4328-a5a9-f06044edce12 · outbound

This paper cites Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.373390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.373390Z digest=sha256:066948667b2c1de64415928d0a0461be4b1e3c30c8b8a79ed5f63766fc4f65e0

Observation 282cb1db-1d9d-4e8f-86a7-92e62621b0dd · outbound

This paper cites CapsFusion: Rethinking Image-Text Data at Scale.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapsFusion: Rethinking Image-Text Data at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.450760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.450760Z digest=sha256:f4053b5d8765c3514ceda6875bbbbbb56bf60ac87bbfbe33cbe9198d33f7608d

Observation b41ccc6b-4fa4-493f-9796-3b6bb1bd2b2f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.594375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.594375Z digest=sha256:ea8c2a2b3ec4b74f7f0589aaadae4769932c981ef2096c2fa85979f378106b24

Observation a5967b2d-3470-487a-9840-053c41c5eff8 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.677298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.677298Z digest=sha256:689ba155fafac23c3a9a38bc812ef94102486117bbd7561d1fc1a3ab3a8b9c92

Observation 38de8f72-b661-4f69-92fa-c19e8eda71ef · outbound

This paper cites online" 'onlinestring :=.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction online" 'onlinestring :=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.782189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.782189Z digest=sha256:5a1eaebea05243c93d058527a5e8c30fd6fe37f8ec2787bf498b86f59c09f75c

Observation b10c9615-855e-404a-ac3e-97799e4bf9d0 · outbound

This paper cites write newline.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.869281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.869281Z digest=sha256:45a192fa8dd2e3e943b90592daabb7e49d18aa3f362f453f2e1b11c3b6b9f797

Pith citing papers

No inbound Pith citation observations are available.