Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:03:55.551725Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 0 inbound Pith citation observations for arXiv:2607.23235.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T00:03:55.551725Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
57 of 57 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de9bb82a-8b21-466b-896b-75f2732a3774 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Image captioning: Transform- ing objects into words.Advances in neural information processing systems, 32, 2019
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec69123c-8d40-43d7-8397-04aa21157e27 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions A comprehen- sive survey of deep learning for image captioning.ACM Computing Surveys (CsUR), 51(6): 1–36, 2019
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b3f5a94-b55a-4720-98f0-129e4b7a8758 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9cb565e-3ac3-4d85-b09e-16220d3f368f · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Surveying the landscape of image captioning evaluation: A comprehensive taxonomy and novel ensemble method.arXiv e-prints, pages arXiv–2408, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1592e1d7-4a1c-4987-96b0-93cab0612cbe · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Revisit Large-Scale Image-Caption Data in Pre-training Multimodal Foundation Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f67b4f26-d368-4036-8a39-a626c20d35d1 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Image Captioning Evaluation in the Age of Multimodal LLMs: Challenges and Future Perspectives
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bacc46c-d093-4197-94d7-7e97b0c39f0e · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Show and tell: A neural image caption generator
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edc2adb1-0f1c-4218-a72e-a33ea1749f93 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions End-to-end transformer based model for image captioning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06e431af-6471-48f8-b589-2aa773f0d7a0 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca182a09-b1f8-4bc1-ba63-74f61bfd27d5 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Cccaption: Dual-reward reinforcement learning for complete and correct image captioning.arXiv preprint arXiv:2602.21655, 2026
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1c1aa8d-138b-46ab-9229-c9bfa5ca1e76 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Learning transferable visual models from natural language supervision
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1924cb55-3517-4cc0-93de-8c1ca9103f66 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Blip: Bootstrapping language- image pre-training for unified vision-language understanding and generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d61b1ac-53bd-400b-8b04-be6af738f114 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc2ebdd-6316-4446-bbdc-66241cca246c · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen2.5-VL Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab0b9b6a-df27-4efd-a05d-bba5c842678c · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen3 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6222624d-ce3d-41f3-853c-e876c5b631ce · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916, 2023
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e255aaf-3065-4dcc-95b7-10398ba50bf6 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4053493b-0866-4e49-b640-edcdf83ca423 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcee4a54-64e3-4c9c-8660-633eec882e8c · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Prometheus-vision: Vision-language model as a judge for fine-grained evaluation
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 824807e0-7ee9-4a89-995f-a8dc7096958a · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Caprl: Stimulating dense image caption capabilities via reinforcement learning.arXiv preprint arXiv:2509.22647, 2025
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5601d520-b86d-4d48-8d07-d7162a776818 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Learning transferable visual models from natural language supervision
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f825394-57c9-4c17-8cac-a67cbc569acf · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CLIPScore: A Reference-free Evaluation Metric for Image Captioning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed5e168c-6f53-4ece-bef1-d084656422e4 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Positive- augmented contrastive learning for image and video captioning evaluation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c275ef-fe7a-421c-957a-f7f46703118c · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Evaluating Image Caption via Cycle-consistent Text-to-Image Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86224940-f088-419c-b747-c63cdb19c656 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Im- age2text2image: A novel framework for label-free evaluation of image-to-text generation with text-to-image diffusion models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d760d816-c7ee-4a39-b9eb-39787985c465 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Visual fact checker: Enabling high-fidelity detailed caption generation
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b557d234-f48b-4829-8f35-10d73131098d · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4454d93-4993-4a2a-ad70-aca87f72da10 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a06b21d3-1251-4760-9ed6-0a91ad2de85a · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Microsoft coco: Common objects in context
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2459e5d-a224-4bd9-8fbf-24c56d8981de · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Vqa: Visual question answering
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea811959-84c1-4db9-9c08-3fa46ba4dd24 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Beyond quantity: Distribution- aware labeling for visual grounding.arXiv preprint arXiv:2505.24372, 2025
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdce35e3-d0d8-497e-811f-e4419971087c · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Mmbench: Is your multi-modal model an all-around player? InEuropean Conference on Computer Vision, pages 216–233
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad42cf7-0956-48a8-a8c7-940d7755b153 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Are we on the right way for evaluating large vision- language models?Advances in Neural Information Processing Systems, 37:27056–27087, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4dab1268-75c8-4190-adef-9f4ae9020ae4 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab00de5-cb71-4dbc-8162-53162fca8cd4 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c6675f83-2f1f-48b5-89c5-2e52fc679dcb · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a2ed12-8357-4d50-acde-0bb217817342 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InEuropean Conference on Computer Vision, pages 169–186
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e5f531b-3f8f-4612-9638-7a2ff6a6df15 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Measuring multimodal mathematical reasoning with math-vision dataset
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da2fef2f-9f97-468f-84a1-07e11fc165c8 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102, 2024
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 636f817c-2c60-4728-be89-ab62bbac3110 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Ocr-vqa: Visual question answering by reading text in images
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd9693b-fd8f-4670-be20-27f3e3eac1a1 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Referitgame: Referring to objects in photographs of natural scenes
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6d4ddd4-45b7-4beb-a220-523c16d3a8bd · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Improved baselines with visual instruction tuning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61ac7eb2-943a-4d9b-acc2-967d7f0b56b4 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Llama-3.2-11b-vision – multimodal large language model (text + image → text)
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4e25302-67e0-48c7-ae84-8151916f4a1e · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6399c4b2-7c03-46d9-8ed9-d68d57647d12 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Finite mixture models.Annual review of statistics and its application, 6(1):355–378, 2019
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ebe9983-979c-4dad-8819-c3f9b26d089d · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions A density-based algorithm for discovering clusters in large spatial databases with noise
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09792c1d-2976-4d88-a576-7955c28e6100 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Qwen-Image Technical Report
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f010210b-df38-4485-9d34-cd89dcb7f9fc · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Unitbox: An advanced object detection network
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62414ae0-c443-4f4c-869b-226b1ea7010e · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Opensearch-ai / ops-mm-embedding-v1-7b
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2606616d-bfa6-47bc-ad81-aa284a0ae4df · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Gpt-5: A new era in language models.OpenAI Technical Report, 2025
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86dd6d9b-d379-489b-b898-047473b38ff3 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Scaling Laws for Neural Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5181199c-25e9-4ae8-9847-8d597757ba8d · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Prism: A framework for decoupling and assessing the capabilities of vlms.Advances in Neural Information Processing Systems, 37:111863–111898, 2024
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc17697f-a386-4aa6-84c2-6440d632210d · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions High- resolution image synthesis with latent diffusion models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7066cee-b6ef-45f8-8ae7-29fb0fed6d7a · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions question
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26925ec0-cdb4-40ee-af5f-f3595014d1df · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Unresolved cited work
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 472851b9-bdb6-470e-9596-53bf21ca11e9 · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions gold standard
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fc606de-a5fc-4260-bbec-4aca9748345b · outbound
A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions Compares the performance of two judgers, Qwen2.5 (Qwen2.5-VL-3B) and Qwen3 (Qwen3-VL-8B)
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.