Pith. sign in

Paper Citation Record · LEDGER

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 4 inbound Pith citation observations for arXiv:2412.15484.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15484 v4

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:26:27.777304Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T15:26:18.843059Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:17:29.017556Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 13794f31-db06-4959-89af-35ff75a8e765 · outbound

This paper cites GPT-4 Technical Report.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.608553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.608553Z digest=sha256:a5c105a0ca31a8f9a62d930079d66641c9dabf11fdc8da68636eb5d8997d5785

Observation d969e1f0-d713-4961-b08b-cf88a20f6cb0 · outbound

This paper cites Llama 3 model card.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Llama 3 model card

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.611236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.611236Z digest=sha256:cbe2309e4c936cf57ccea1035d7a0b220e422fb299418c2b2342eb7675ec312a

Observation e73edd62-b464-4876-a4e2-0ae2301dd801 · outbound

This paper cites and Lavie, A.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage and Lavie, A

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.615692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.615692Z digest=sha256:f66e23e9c6f78ee71fb1b8685a3853a69d0467ba58f7b96e11f2ab0a68c8fa74

Observation 26711d55-6b4b-4135-adf8-716fb3a0a8b2 · outbound

This paper cites Language Models are Few-Shot Learners.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Language Models are Few-Shot Learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.619262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.619262Z digest=sha256:930f0de632591499ef0fd36d4021c59f71c3a09b6456dec799b1d71f14387cd6

Observation 18c03a80-2577-415f-b5f0-2975e3a93ffc · outbound

This paper cites Clair: Evaluating image captions with large language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Clair: Evaluating image captions with large language models

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.171064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.622710Z digest=sha256:a02115b2b9e102cf4b7bf60d9bd497ca9edf4a7669b5a764a3f2acd3d350e1c0

Observation 55d7c217-ef28-41ea-b2ab-fc8f1e74e81e · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.625948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.625948Z digest=sha256:6a0d5fdbbcf72fcfd3de3f4ef919a23c55559f8c2fe5ee12db0d2666fbefdbe7

Observation 910363c6-ac62-44b9-8f6a-8b69a7030d61 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.629875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.629875Z digest=sha256:5d207f0e0f1625a758e86d56d08dc4191b9d95ce041d8f928ba570f4d9038acc

Observation 2713d1f7-686d-4f01-ac45-9abb6baff0f2 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.633432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.633432Z digest=sha256:66a915709cee5f401ba4dba66894d60c8078405790cd1bffef098d797ac987dd

Observation ec94c0ec-f2ca-4375-a7a2-5a634b673983 · outbound

This paper cites Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Efficient and Effective Text Encoding for Chinese LLaMA and Alpaca

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.637305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.637305Z digest=sha256:b94a9229e553206ebcc5293374e126a3b8728557fee671066ae8fd15c7538f28

Observation aaeeac2e-f225-4aa2-a8d8-c5e28025d2c8 · outbound

This paper cites Instruct BLIP : Towards general-purpose vision-language models with instruction tuning.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Instruct BLIP : Towards general-purpose vision-language models with instruction tuning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.641246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.641246Z digest=sha256:66e0910f1a0d373c039f5ca6c84f4e5994307cf063bacfb8bb0371917a9a77ca

Observation 302ada15-5c9f-4496-a6b7-5971401e8e94 · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.644916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.644916Z digest=sha256:d27bede7ade6a337f149f8f0e4f419819feb420748b07fd0c6d0b42e871033a5

Observation ad362396-15af-4ad9-aa7b-314bf0cd21f2 · outbound

This paper cites ImageInWords: Unlocking Hyper-Detailed Image Descriptions.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage ImageInWords: Unlocking Hyper-Detailed Image Descriptions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.647518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.647518Z digest=sha256:f81c6a9910e29beebd0c5e4068209dafa737dd7dacb08f5b16c0e3cf5f65f155

Observation a6caa144-c09d-427e-9444-1adfc074046a · outbound

This paper cites S., Lin, T.-Y., Liu, M.-Y., and Cui, Y.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage S., Lin, T.-Y., Liu, M.-Y., and Cui, Y

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.154229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.649856Z digest=sha256:9e052ca1280d3290472471c07b670084b84af1f2b5eca392ef7af9cb2d248c1b

Observation f6647f82-f181-4143-895e-7fd82a42df78 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.144451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.651866Z digest=sha256:fda2510dd7381b4f1b482aee5f2fda96e931d52717d96119c11a3249e928bf1b

Observation 329f3145-8682-4acf-abba-751820cf40b8 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Clipscore: A reference-free evaluation metric for image captioning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.135998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.654346Z digest=sha256:fff0e018fc0d9e8022693d9d6088e7f8e34d1fb7962f500b21b4e1d156812608

Observation ac6c0219-a870-4e5f-88b8-671f20d212ef · outbound

This paper cites Z., Sohel, F., Shiratuddin, M.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Z., Sohel, F., Shiratuddin, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.128179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.656602Z digest=sha256:48f9718d2bbda0e62e91a66ac860b668265b363e020670297a6bc0b28abf2505

Observation f1558ded-ab1d-432b-8f30-3204fcf561c5 · outbound

This paper cites Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Opera: Alleviating hallucination in multi-modal large language models via over-trust penalty and retrospection-allocation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.658517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.658517Z digest=sha256:a5fcf02171a39f80ca32b8def14221f187e84d7132e608d356098c53b303c3ae

Observation 20741046-4c54-4e5e-a3db-757ed355ff83 · outbound

This paper cites F aith S core: Fine-grained evaluations of hallucinations in large vision-language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage F aith S core: Fine-grained evaluations of hallucinations in large vision-language models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.660944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.660944Z digest=sha256:310837bea1261b4b03ffd0b6e7256a87635aab0d5c49f8f7e85aa91b83f22add

Observation 92c03982-9592-4fa0-8566-8fba264fa55d · outbound

This paper cites Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.663250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.663250Z digest=sha256:8fed738ca6ecd6c06345b921cee51ab48967efa3d7ae526dc2b9baaa826900d9

Observation 52533384-95f4-4dfa-8860-aa345239dce9 · outbound

This paper cites A diagram is worth a dozen images.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage A diagram is worth a dozen images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.665277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.665277Z digest=sha256:9bf0bea575e4bdb1c347fec0c5a47601dba9742bd4ba796d8931bbaeba691f6e

Observation 0aabc5b8-e974-4b39-9c2a-74adb0a9912a · outbound

This paper cites What matters when building vision-language models?.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage What matters when building vision-language models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.668185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.668185Z digest=sha256:bbd74b45ce296e26a5c3a9c6fa7820daa59ea955cc5037867312c98d7d5cb319

Observation daa49e4b-9ee7-4f9c-ab74-7689ce256dbc · outbound

This paper cites Volcano: Mitigating multimodal hallucination through self-feedback guided revision.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Volcano: Mitigating multimodal hallucination through self-feedback guided revision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.112048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.671738Z digest=sha256:2b267003fc8efa58c78bdb4181a09ad7d29bf637545cbf528dc55340a454e875

Observation 6d068fd7-f495-4e9e-be4a-5e10c8d9817f · outbound

This paper cites Mitigating object hallucinations in large vision-language models through visual contrastive decoding.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Mitigating object hallucinations in large vision-language models through visual contrastive decoding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.674998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.674998Z digest=sha256:c610287f83bbac88cf88cb6520711a92e62d47fb8a810b955fd9f4b4c59991c1

Observation 4627dc9d-bf33-4318-8ba4-82caf8dd00c4 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.678292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.678292Z digest=sha256:7d36e7ec2522e8fb0983128f558ebd93edf428abbc85afead8a7795fa0b73533

Observation d5c38698-9bbe-492e-bf47-c751068b0acb · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Evaluating Object Hallucination in Large Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.681197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.681197Z digest=sha256:25b772dd46b800c3b3d9c0f18f4e57345d778287c315c613167376611813ef31

Observation 0d420963-913e-4eb1-86b9-dbfb51cbeb37 · outbound

This paper cites Rouge: A package for automatic evaluation of summaries.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Rouge: A package for automatic evaluation of summaries

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.684274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.684274Z digest=sha256:02f0d05c678101849cf115f74154dae75a1780affe85b4ebd7b6395c328f7e1b

Observation 1811c243-dd95-4ea8-b8a7-d9b7d21568de · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.089954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.687351Z digest=sha256:ee20cdb1cfc01a136c1c9b24020754053a2cd847708f49952f935d6ccc33ac73

Observation 48d4fff8-92fd-4cc0-b216-5b506e9ad961 · outbound

This paper cites an unresolved cited work.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.690644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.690644Z digest=sha256:dd6495909bba0ef70e1abc194e992b94fc46b0bee871bec730846e69c77c37e4

Observation cb62b5ea-5732-40ae-b808-9d33677e949f · outbound

This paper cites an unresolved cited work.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.693997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.693997Z digest=sha256:442ecd0e372663ff593c844668d05ef83871d25ee820da9314bc692d08ce4bbf

Observation e44469eb-48c8-4114-9a4a-99bbf5cf0061 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.697443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.697443Z digest=sha256:8b63550279abc6dead083b1ce29f71fee03a1de81cda0a394f2a08cc6932072b

Observation 31ea0a97-795a-4426-92f7-036f022cb01d · outbound

This paper cites Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Paying More Attention to Image: A Training-Free Method for Alleviating Hallucination in LVLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.699612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.699612Z digest=sha256:41e750544415707f5cc3a3edebe1e22a5f630832a05c561b0d2166115ece9271

Observation 808d8ad8-a3a0-49cb-9fdd-4f89bdd04a7c · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage MMBench: Is Your Multi-modal Model an All-around Player?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.701965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.701965Z digest=sha256:d373beccabbdb7233f81102d29b6925755b109c16b522d9681b269cc4ad9b2aa

Observation d5549d7e-0a58-48bd-8540-292031b5e641 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.704419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.704419Z digest=sha256:ae273ce8fb08977345dbd41425226f5157a08f576cc4b2bfffce6782a5062785

Observation 87ead6b9-409d-4014-9406-f456ea424c87 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.706631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.706631Z digest=sha256:9331932cd168c52a56ba1d8b3d9b818342b79f103d15185d1fb676e17199a149

Observation d7716b99-e09f-4546-ae43-fdd5d175f567 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Self-refine: Iterative refinement with self-feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.708864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.708864Z digest=sha256:5a3f308505e41b0fdeb75543832d73d3561f2a6e9a11638ba662b050461c8c1b

Observation 426f86f9-e2a0-4d1c-8f5f-d325d5f61f67 · outbound

This paper cites On faithfulness and factuality in abstractive summarization.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage On faithfulness and factuality in abstractive summarization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.064977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.710842Z digest=sha256:b63e5488e17ce6cef997f1af0ceee38187eb664c2a930888543f2e266c7d3f00

Observation 8f1cdc2d-b278-4b40-8f8b-8145dcd3b25b · outbound

This paper cites FA ct S core: Fine-grained atomic evaluation of factual precision in long form text generation.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage FA ct S core: Fine-grained atomic evaluation of factual precision in long form text generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.714096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.714096Z digest=sha256:782f9d2aba3c39775bbd4851b101444951971fd0187c0f5f72c5a434b9d17b55

Observation 796b09f0-3f6d-4d33-ad5e-6fdca5defe66 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage DOCCI: Descriptions of Connected and Contrasting Images

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.718006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.718006Z digest=sha256:4ea7d44ea77f9758f951d84c7b5ccb69c04710e0c108720c248a8cf0308e037e

Observation 348fff99-694a-4a6b-8378-99980841983a · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Bleu: a method for automatic evaluation of machine translation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.721620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.721620Z digest=sha256:aabd81f30ba052e2ebd8f925ab98397b94ead331dcfe45ac2f02842a159fd8a6

Observation 48b37bfd-b999-4588-a844-5ac702a6fd5c · outbound

This paper cites Aloha: A new measure for hallucination in captioning models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Aloha: A new measure for hallucination in captioning models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.053336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.724581Z digest=sha256:fd4adf2cd8ea49199d7f68fef6f6077bc1207bd37136cf267bdd14678ff8c3a7

Observation d3ba5887-53d2-4d5c-a7b9-9d94fb8487dc · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.727345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.727345Z digest=sha256:0cb48aa221dd1116a05c318164b69bb3c3af0c1de06c83bb84fca1a7d14003c4

Observation 2cf316a9-ac13-45e4-b22b-37611da7dca9 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.730133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.730133Z digest=sha256:4db4c43bb86a6422ad41aa90f6527a859cb67ea263499a7a9f5ffd16bc4618fc

Observation f8176436-42d6-401e-b2b8-44947dc1a70f · outbound

This paper cites A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage A picture is worth more than 77 text tokens: Evaluating clip-style models on dense captions

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.041269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.732745Z digest=sha256:15804ea9972f522e74417a52f563cede679dfc4cab419ed9e4de9d35caba4c26

Observation 10dc727c-3e31-4465-b4f1-e174d057439b · outbound

This paper cites Attention is all you need.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Attention is all you need

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.735477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.735477Z digest=sha256:a464373e98e3f6db272b7ecbdac6722f6017317ef2ca26b09c6774e6e4aa33cd

Observation 46b86ecd-1b2e-4005-98e2-bc7bcf3b67c0 · outbound

This paper cites Cider: Consensus-based image description evaluation.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Cider: Consensus-based image description evaluation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.737938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.737938Z digest=sha256:b4106af64dfc899e4e0f68e2ae8df316caadfcf361329aa6d250444452b88512

Observation ddf002b5-609a-4b0c-ba6d-c4c47f2400c9 · outbound

This paper cites Show and tell: A neural image caption generator.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Show and tell: A neural image caption generator

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.739874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.739874Z digest=sha256:1faa5fff13f710a632b6f6e58a9b373a8ae6f28c711078c1ca1181779202b686

Observation 8b383e42-4b22-44ed-b7d6-d009c41d38d5 · outbound

This paper cites V., Chi, E.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage V., Chi, E

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.019824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.741863Z digest=sha256:f0539d4d71792cfc4cba2ea2eaed48c44e4a4cb37677b108081bcb166cdbfe23

Observation 5304c74a-5b2b-4eb0-836d-ac3a3730c39f · outbound

This paper cites Show, attend and tell: Neural image caption generation with visual attention.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Show, attend and tell: Neural image caption generation with visual attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.743859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.743859Z digest=sha256:60f209f63f4d01f3484cc2da89be6439acc2bc8295a084d89d0b4457b9aedbd4

Observation 42064b2c-5b10-4d13-aa5e-460b5cb6e895 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.745848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.745848Z digest=sha256:3b949fa8d0b6157140dbe0e95b052b57f562c05412fca73bcf7c9a5306d18bf9

Observation 517aa297-0f10-4a82-8133-ac87e6ff4c73 · outbound

This paper cites A Survey on Multimodal Large Language Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage A Survey on Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.751320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.751320Z digest=sha256:d5e1865bdc430d3ad2043f7490eba7e6dc8f585136c90602370b8d864d30cc17

Observation d13f4ca5-a833-4186-b14e-b149d8764086 · outbound

This paper cites Woodpecker: Hallucination Correction for Multimodal Large Language Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Woodpecker: Hallucination Correction for Multimodal Large Language Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.754577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.754577Z digest=sha256:6dcb30198947cf98a0995a98dce2ce9addee720d3c7455f977be5ae14fc80176

Observation ba5721c3-85d1-45dd-85f7-38222a4d27b1 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.757764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.757764Z digest=sha256:b4e4f24b22915be619e2898b24df8547211a40943a1b78a573283a54f0b7171d

Observation a14ba7c9-619a-4b42-9742-91fb24817973 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:28.006318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.761657Z digest=sha256:2c2fc9510c27c795a0c8f37db15ae71096418efad134075c2fc6164d3dccf409

Observation 8b4bbfe6-dcf3-40d7-9be7-96257f5ad39b · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage When and why vision-language models behave like bags-of-words, and what to do about it? In The Eleventh International Conference on Learning Representations, 2023

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:27.998566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.764401Z digest=sha256:406fe383f2036ceee1584fdc79673b92bf59e793f485dd92fe5d0dfa2c8a2c6c

Observation 9683cd08-6e6c-4b4a-bcb6-12263bf24d39 · outbound

This paper cites Enhancing uncertainty-based hallucination detection with stronger focus.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Enhancing uncertainty-based hallucination detection with stronger focus

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:27.990593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.766917Z digest=sha256:94a2cf5c664955988621d148210da58575d759f22084fb69949878830c442a3d

Observation b36e7235-d537-4ede-8b5d-2186e8e22a11 · outbound

This paper cites Knowing what llms do not know: A simple yet effective self-detection method.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Knowing what llms do not know: A simple yet effective self-detection method

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:27.982195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.769110Z digest=sha256:081d9b4648a113da604c8ff981367c25ed1c97b586db7bc115b1fdcc929a5ee0

Observation c3239953-f57d-4c13-8a3f-2a976e470b8d · outbound

This paper cites Investigating and mitigating the multimodal hallucination snowballing in large vision-language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Investigating and mitigating the multimodal hallucination snowballing in large vision-language models

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:27.972116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.771052Z digest=sha256:35ae938656f439ad2729cbca475691a76ea1ec40a753f63304abbb6a97f637d1

Observation 17f7c056-1188-430d-81e4-84d2c5403574 · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage Analyzing and mitigating object hallucination in large vision-language models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T11:26:27.963823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-11T11:26:27.773198Z digest=sha256:f313890a70e96d4fcf728ddc646e40a59cf385b9c1397ad20860ad81abed5896

Observation 028ebee8-01d2-45f4-82d5-73e1be346cef · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.775253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.775253Z digest=sha256:b7dc035bec0148698b054f7e0ce447df0e9e5f8b88566a4cd94fef69f9bacf40

Observation abb8c248-ab28-4181-a59c-5c85b5baf2a5 · outbound

This paper cites write newline.

Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage write newline

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T11:26:27.777304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:26:27.777304Z digest=sha256:ccdb2c837cdefa6192c1d4b07b2dfccb5c072971ea0db349f304adad3e45430a

Pith citing papers

Observation 4984a3af-918b-41e9-a7cf-d6276e863dae · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.843059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.843059Z digest=sha256:973718524e92a196e5396bf639edb999fabb10a7ccd0742f5c4aebdecf39068e

Observation d0338b92-ab70-4b33-84f3-8f52ba425fb5 · inbound

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation cites this paper.

Steer Where It Matters: Token-Level Visual-Sensitivity Steering for LVLMs Hallucination Mitigation Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.642228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T10:22:40.055153Z digest=sha256:f3528af3326bdfdc27c05878df3453bd63d965a9729f3bd25f03428253ac9f1d

Observation 1c122c86-07e9-45f2-81ff-a39571f7cca3 · inbound

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning cites this paper.

CapRL++: Unified Reinforcement Learning with Verifiable Rewards for Dense Image and Video Captioning Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.019126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T17:21:38.543724Z digest=sha256:124a4701d2ce76f6105e05978204045806c45ff0714568c89748da691ebb43e3

Observation 893d933d-6f00-4da1-84f3-d36db0bd92c8 · inbound

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation cites this paper.

CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Toward Robust Hyper-Detailed Image Captioning: A Multiagent Approach and Dual Evaluation Metrics for Factuality and Coverage

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T04:09:50.687652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:09:50.687652Z digest=sha256:2053e70fe413b5cce4b2547853bd2886077fecfda7c39941b3678190a05acc3f