Pith. sign in

Paper Citation Record · LEDGER

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction

As of 23 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2505.22613.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22613 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:08.869281Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 90463c61-bf5e-443d-8851-0b8611e30f46 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.224914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.224914Z digest=sha256:2f2b8c1685b507338a4814850db81170d6375d3fc42b513482f48652a7803b05

Observation dc5ff3fa-1c18-47dd-b333-692732d69e67 · outbound

This paper cites SPICE: Semantic Propositional Image Caption Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction SPICE: Semantic Propositional Image Caption Evaluation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.296549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.296549Z digest=sha256:ea3133904638c431fc6b6b858f8adbf6eb4d8ef2a51d6534e5a0c2e805e1dba1

Observation 714fbf1d-f07e-4025-aea3-127cae2dfbc4 · outbound

This paper cites Qwen Technical Report.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.411249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.411249Z digest=sha256:353038b481be1a5ed86524891ea54ba3455ff30416a601cc689b3edbf52e57f4

Observation dc6bcb00-ea6c-45ef-a8e1-9e2c5364cc79 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.514104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.514104Z digest=sha256:de81fcf1d30ec058ba88ad2b8cef4cb27087749e0008bf3e73425ca84bb31a72

Observation e803da03-b8c5-45fa-9602-7b9c93be180c · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Hallucination of Multimodal Large Language Models: A Survey

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.609812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.609812Z digest=sha256:528841a2eb023dbb175b0095b7b7c8de5314f10006f5bed680266d22b7257c4f

Observation e97fa7e0-4965-4cb8-8b22-5fc9a85af19b · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:08:10.300178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:08:02.710978Z digest=sha256:b451dc1219273423c5cba13931650499191616b95255734d3c503bb1cc9eb21b

Observation e059b8d5-4d35-4768-99d0-84a0df065497 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.797253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.797253Z digest=sha256:501a73325011d4fdc0ddc6fa63aeb2be7849a8f5fc6393420a04adf98c148372

Observation 5cb60e30-a5c0-4049-9cb6-7d47de177c7f · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gonzalez, Ion Stoica, and Eric P

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:02.914321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:02.914321Z digest=sha256:88577dabd7d524cc5eb595aa40174ba328dab6c447e5366ab1877217ad43e804

Observation ec6b558f-1628-4982-9668-2ca3c7f3a36b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.022462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.022462Z digest=sha256:ce7059c4e592e5a30513961746d3ba96239af2e8ecd4fe3d9cf83bbc1c2fc0e8

Observation 285fe182-3d78-4df4-860c-79512ddbfdd4 · outbound

This paper cites Benchmarking and Improving Detail Image Caption.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking and Improving Detail Image Caption

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.153714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.153714Z digest=sha256:c01f2b20dfb28941b91043f9f67156ed83926e6260c5f889c8dae0e15b454263

Observation 93f5842e-428b-43d3-8fb6-f97863357a38 · outbound

This paper cites A Survey on In-context Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction A Survey on In-context Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.233342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.233342Z digest=sha256:b738918e926e21c7d583aa007fedddd9b31a6af1e9455609988a6039d32a9b03

Observation a0697e11-d0e0-4a02-872b-e8111bb91916 · outbound

This paper cites Improving CLIP Training with Language Rewrites.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Improving CLIP Training with Language Rewrites

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.303022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.303022Z digest=sha256:5952ea5978f20a76fb31a37612d58d8470204abeec803aba0b18dc98f45f9674

Observation e1c9293f-1173-4e62-b229-6156df03b9d9 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.374847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.374847Z digest=sha256:e4da8ff17780d981ff4a8d081877cc02d0ea8fb62cd43082f39a7cc5a8363525

Observation e9b1095e-54f8-46ee-9570-65993b2f0c2d · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.447542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.447542Z digest=sha256:1f5ddc6bc65d1e6cbfb90c730b30b60b2a65a2e5ecbdc312816e94435eac0d70

Observation f17a7886-2c4c-4fc3-b881-22885e2569bc · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.524335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.524335Z digest=sha256:a1ffc1f798975af426c299a6fb39760d5b06f4d07fe8a20a28a79674bc4ccc0f

Observation a683d7db-d212-4e2f-92ee-460c5bb7e6bc · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM2: Visual Language Models for Image and Video Understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.626178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.626178Z digest=sha256:92740e3420ce9f8956573eff290a06ad7ed28bc5d5538b40274b916a44c4d9c4

Observation 8e721d09-db49-4d03-ad35-f21644d1adc5 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LoRA: Low-Rank Adaptation of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.741011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.741011Z digest=sha256:31331e8d1bfbf354329318b9df9cfbbd6a7c46a3eab26f39292b076521f97d8b

Observation 67873c65-c2b7-47b4-93ef-2441ad0781a2 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.841863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.841863Z digest=sha256:3f1b6307e122302119b6d75493b40486a0d08498f9449f8c97c02f0a96b68741

Observation ab514049-398b-44d1-9f7b-4630afa36023 · outbound

This paper cites T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction T2I-CompBench++: An Enhanced and Comprehensive Benchmark for Compositional Text-to-image Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.913503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.913503Z digest=sha256:f74b9ec4a2042305091f8425d6dca7e0fdb67d781e62e7fc730255fde96c225a

Observation 65f33dcf-9ea4-497b-a508-52d258668a74 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:03.998066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:03.998066Z digest=sha256:a9a5b261b39c7b82bb36660b4689534f8cde8e89855c0e5610a846a64e862b73

Observation 9e7fa958-1490-4002-81da-fac08bbf48bf · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.091130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.091130Z digest=sha256:285141a7b2bc5c9cdb62e2d9b6d637cfea6f93bf18bee0fbd88fefa0e7d9015e

Observation ae4bd084-619c-4959-a82d-421bd37155db · outbound

This paper cites VeCLIP: Improving CLIP Training via Visual-enriched Captions.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction VeCLIP: Improving CLIP Training via Visual-enriched Captions

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.205365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.205365Z digest=sha256:eacd2120eecdffaaa2998815590c4e3040dbd061526f012aa60154a5acc8d6aa

Observation 4af59865-4ae8-42ba-a9c6-40fe3663865b · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.303127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.303127Z digest=sha256:1277d3659925f6fde07acd6e4e6054a0eb072e1a879dcbf044e93474c78ed691

Observation a14b41b1-eede-4c43-b21c-bf8d2717c0e5 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.386112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.386112Z digest=sha256:96fa3bceed4b92db4f1d1f855f60bc01ea4079a8dfcee9cd6c31da8c14b4cb26

Observation 50aa06cf-5d1b-49a6-8515-95435e9fd6d2 · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.479059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.479059Z digest=sha256:1288a27211b8eb00bd608b339c199b3b8f4c4fb426f953e06209e1a17e7bf2cf

Observation f82f29cd-725b-4492-89dd-b8a0c02d954b · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.565373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.565373Z digest=sha256:d04f5cb7301bf3a0b9f554842a0e7aae610d55125363891cf5e2255cdd366251

Observation f79b9c52-9e3f-46e8-9e9a-96587df0c713 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.662605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.662605Z digest=sha256:43533c7230b668685caf87f241b321fb06d54050458786332afb2b271887527d

Observation 25002f22-4f2d-4fbf-9222-f49370bb2c88 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.804019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.804019Z digest=sha256:8a5fd67e360ccfd364c572534c9235120c34b29b828d19b38fb89f88e15b95b7

Observation 27ef73e7-64f7-4ca8-92d9-a7a406674719 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Improved Baselines with Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:04.893325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:04.893325Z digest=sha256:059190b9bfc4c0e189d70f87c48153b11b14b0b23b116e275c0ebe4cd4c73535

Observation f3c8e907-8dcc-4ff2-93ed-c5653997fe54 · outbound

This paper cites Visual Instruction Tuning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Visual Instruction Tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.034601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.034601Z digest=sha256:d399d0e62ab2c3cadd8059d22e6445dff1b87de2a20722474ff337091064aa2f

Observation 35c61650-6da7-4968-a50d-1fe8be2d38ea · outbound

This paper cites Decoupled Weight Decay Regularization.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Decoupled Weight Decay Regularization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.189909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.189909Z digest=sha256:d7d5aa2afa81aaf578f100e6d7c1b41954b5307636ad8c3ea2609c97838c70a7

Observation 050d1fec-5a17-4cbb-a3dd-6661b90d05d4 · outbound

This paper cites Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Benchmarking Large Vision-Language Models via Directed Scene Graph for Comprehensive Image Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.395780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.395780Z digest=sha256:1df0def0c97e81c2dded835513030bcb6a70eb1e07f38ccc928e619220522f38

Observation 7fbd89d8-0707-44bc-b076-3e5942c8d3fb · outbound

This paper cites GPT-4o System Card.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction GPT-4o System Card

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.567642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.567642Z digest=sha256:c834efce636c42725b56abe167e5aef48e11b417cddfae9cc5faa23ff402dd98

Observation e12b594c-8055-4723-88d4-332c0e529590 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.741641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.741641Z digest=sha256:2aaff31e07ca09de7903b77246d6c6752e0ab912e93bc49b617fad333ac8186f

Observation a1f97d57-0233-458c-8d8c-b6f5ad7bcb87 · outbound

This paper cites Training language models to follow instructions with human feedback.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Training language models to follow instructions with human feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:05.893497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:05.893497Z digest=sha256:eff958b2027a6aec465c807b14d17108541d95929d63339d3f8b4da1a1bf322b

Observation 9e3b42fd-d060-407b-961f-0a91638baa29 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.015674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.015674Z digest=sha256:b395cf9360a83bc541836d2f0d6c702fe403c759d2fd4ec252e35ce1bc31d0dd

Observation f8243aab-30e4-4d78-8449-60e07b5b33d0 · outbound

This paper cites Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Patch Matters: Training-free Fine-grained Image Caption Enhancement via Local Perception

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.657590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:08:06.136121Z digest=sha256:3712c9ada221bfebd1d4f9d46ddb6ebda814d97673452a12655f50805b1cfcd1

Observation b96e1fe6-ade3-4e55-99f6-14941409d313 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.252023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.252023Z digest=sha256:16a8fc293a0b449a644f04a27d25d04c11a472f937cc7cb4d423150cc193108c

Observation db01e3cc-8589-40c3-811e-42278e299bee · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.370126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.370126Z digest=sha256:3a8af8230a3005383f1c9e8038492c3c1028db78ca723fa91e24e9226d8ee543

Observation 0bbf7527-2e0a-43dd-8862-846ec2812208 · outbound

This paper cites TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction TimeChat: A Time-sensitive Multimodal Large Language Model for Long Video Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.501876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.501876Z digest=sha256:4836226ffceee087401a7f66408b61b22e092064cea46ee83f2bde0db8b0688d

Observation b91d8922-fafe-4c54-8071-bc1d40044a68 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.641934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.641934Z digest=sha256:6ef764f8d4c6e283a55b2d9dc297c4e14cd727b60ab78d6c9e5ed869655ebd99

Observation a85e9afc-a2df-4d8a-9b1a-719b8539f95b · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gemini: A Family of Highly Capable Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.759252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.759252Z digest=sha256:a81bb12d6bd418340e2ec721827e11063bd55b278751d73d94d54b079003efc6

Observation 24b18474-4d6c-4d2b-a31f-3b23fe409bb8 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:06.905429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:06.905429Z digest=sha256:8586208992ff865c79d5d82fd08461631a56c995a65943d7a7d9bd38cd37dcc9

Observation 1fe5ad3a-fbac-417a-9cd5-dfb1c964319b · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.021522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.021522Z digest=sha256:5c51930da76341d38786241cfc3cda19ad275a04d9454d5aac9d24fbba17f93e

Observation 1f58cc58-b22b-45a3-ad0c-4314b2fb9870 · outbound

This paper cites CIDEr: Consensus-based Image Description Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CIDEr: Consensus-based Image Description Evaluation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.125212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.125212Z digest=sha256:23e5ef1f0451409e20f854fd8aba561efc2a7e8941ef7c959a73de01e8a17fa1

Observation 30f5fc9a-3a27-4395-af93-e472c428b700 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.239579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.239579Z digest=sha256:6602657ab22061eac93141eef590683f33d7475e8907bef8d33693c5c21e9949

Observation da252b43-a893-4478-abe5-2eb5a55d4d6e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.365190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.365190Z digest=sha256:83c0a299600ee5246f90ce5920a0872bd1cf55fd54af500a89f433e6678f9171

Observation 10e2b154-614f-429e-bd7e-aceb10d84581 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CogVLM: Visual Expert for Pretrained Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.562423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.562423Z digest=sha256:4e77ff5421215429cbbbb4d9ce0711862b607bc5f8686a3d4734d9d236e2f4b3

Observation 0946a2de-47c2-4566-92ca-7f82c54c7637 · outbound

This paper cites LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LaDiC: Are Diffusion Models Really Inferior to Autoregressive Counterparts for Image-to-Text Generation?

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.404505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:08:07.687577Z digest=sha256:82b4252bb4e346836f5d3d8071e58a78d0f2a2ea16aecf8852b986a4e51af33c

Observation c8d4c1e2-bb6e-428d-984e-240281321dc1 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.783271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.783271Z digest=sha256:07402ef26e0b3089e42d0202234baf23d95fae062450073dfdad1acdafc2c42b

Observation b48f0b6e-904e-436b-ace1-e253840de87c · outbound

This paper cites Altogether: Image Captioning via Re-aligning Alt-text.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Altogether: Image Captioning via Re-aligning Alt-text

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:07.915717Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:07.915717Z digest=sha256:340987abed1ae0b17f32f0486720c18d1e2dd0f3f433dbe915ffb52c83c310fd

Observation 4a1361b2-4b68-4d2e-adf4-9a1239c4f5a1 · outbound

This paper cites an unresolved cited work.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Unresolved cited work

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.086271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.086271Z digest=sha256:32e03ea4e3d3e39f0118d3dc67d9fb6c96620536655d488523b43d83fb7d36ea

Observation e07b8e6d-61fa-4e78-9970-982a6a4e66c1 · outbound

This paper cites CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapEnrich: Enriching Caption Semantics for Web Images via Cross-modal Pre-trained Knowledge

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:08:09.129159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T13:08:08.229404Z digest=sha256:ce7c24a2925c0a8fae9da5da3b0b6a6d27a96385cb7bd09135bf88254a9af236

Observation 264c1c5c-0e8a-4328-a5a9-f06044edce12 · outbound

This paper cites Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction Painting with Words: Elevating Detailed Image Captioning with Benchmark and Alignment Learning

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.373390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.373390Z digest=sha256:aaaa6f31094455c8c32386ffb325985d4d71e696d59ed3e76142f197730ca16f

Observation 282cb1db-1d9d-4e8f-86a7-92e62621b0dd · outbound

This paper cites CapsFusion: Rethinking Image-Text Data at Scale.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction CapsFusion: Rethinking Image-Text Data at Scale

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.450760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.450760Z digest=sha256:365e4c2e271189c084030267f65cd20f21bebe2d14ecb68a71eb0e53cd7b1e6f

Observation b41ccc6b-4fa4-493f-9796-3b6bb1bd2b2f · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.594375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.594375Z digest=sha256:3de38373aa2bb28b37042a046dabe3b86af5c104dc2a2e4e2fafd2de87671cc2

Observation a5967b2d-3470-487a-9840-053c41c5eff8 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.677298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.677298Z digest=sha256:414c94a990ebd9242068861f69b692a62db042927fe9995afb397367464d67ca

Observation 38de8f72-b661-4f69-92fa-c19e8eda71ef · outbound

This paper cites online" 'onlinestring :=.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction online" 'onlinestring :=

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.782189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.782189Z digest=sha256:cd03d9442ade7ea3995887d834f01b0cef9a57102de98d4ffb7339cc24c701d9

Observation b10c9615-855e-404a-ac3e-97799e4bf9d0 · outbound

This paper cites write newline.

RICO: Improving Accuracy and Completeness in Image Recaptioning via Visual Reconstruction write newline

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:08.869281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:08:08.869281Z digest=sha256:0c14afff293b685b461779bc8f29d9392c1e13b82077b7de6f426827d65553ad

Pith citing papers

No inbound Pith citation observations are available.