Pith. sign in

Paper Citation Record · LEDGER

CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2503.12329.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.12329 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T23:42:35.033889Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T12:59:53.203155Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0c5da5d0-b350-4d61-90a5-265e4b549417 · inbound

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework cites this paper.

Towards Scalable SOAP Note Generation: A Weakly Supervised Multimodal Framework CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:35:41.711143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:35:41.711143Z digest=sha256:e0f8d2363f411949f904b51b980dfd45dd1abee3d2b4e6065b8e5bc81abb92fc

Observation d7974b22-fb85-43b9-b84e-fead96685d62 · inbound

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes cites this paper.

Skin-SOAP: A Weakly Supervised Framework for Generating Structured SOAP Notes CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T23:40:43.812735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:40:43.812735Z digest=sha256:ed16179f5faf73cd05560c5897124d30c86e30cbd3696f368f0cf46a88bcb418

Observation dbf31a79-d4f5-4590-a3bf-91095cc033ad · inbound

CaptionQA: Is Your Caption as Useful as the Image Itself? cites this paper.

CaptionQA: Is Your Caption as Useful as the Image Itself? CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:44:07.225915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T05:42:57.288054Z digest=sha256:a12cc48151b31ff785d7324666d42b895ca0f386b227589034a5c2171d461a18

Observation 574ce0d5-7389-4f12-ae23-a6af18a398a0 · inbound

Vision-aligned Latent Reasoning for Multi-modal Large Language Model cites this paper.

Vision-aligned Latent Reasoning for Multi-modal Large Language Model CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:50:44.238024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:48:23.272002Z digest=sha256:d4651e326bc15ac1f9bdb248894cfea2cfa50c36914cd07262cd5b1fb8b578d0

Observation c230eead-ff46-483a-99ae-30ad6364a4d3 · inbound

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation cites this paper.

MSD-Score: Multi-Scale Distributional Scoring for Reference-Free Image Caption Evaluation CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:46:07.552941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:10:28.941565Z digest=sha256:ddc7cade4a8b66e2a69d2e29d5f712ad670c39ec74609d2faf423edefa5e84c6

Observation 9f292d24-b66f-4b10-95f7-9c31d2ab9caa · inbound

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning cites this paper.

BalCapRL: A Balanced Framework for RL-Based MLLM Image Captioning CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:15:58.030863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T01:52:27.764121Z digest=sha256:b9299e00b406d7b27425fbdbc8208f5ddd878d6844bac23d5b8b6d596f2c9196

Observation 21347e65-6998-4b47-b65a-2b7dec7cc69e · inbound

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data cites this paper.

Context-Aware Spear Phishing: Generative AI-Enabled Attacks Against Individuals via Public Social Media Data CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:52:05.356803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:48:46.380615Z digest=sha256:0d9e476ec382cd7a28985b0c539284de7e682f29ee1bc446a1fa086196d5b3a7

Observation d10a42fa-dc51-4dee-af00-fd5b4db9b0f0 · inbound

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models cites this paper.

Seeing Time: Benchmarking Chronological Reasoning and Shortcut Biases in Vision-Language Models CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T13:06:59.633495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T01:30:36.491137Z digest=sha256:bd63034c0f1c8b3bea00142fb770e4318f37ee5e828d48a60d1a93dc26d61e60

Observation fee9f304-a5db-4605-aa0a-b1b01eb3acef · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:59:53.204581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T05:33:25.373918Z digest=sha256:7fed38d8b57c558e56c2ee06a5b161a9ae4e7d15ae61e50655242b111670f925

Observation 05c19696-2c8d-430e-8cec-d4eb7a2b22f7 · inbound

Robust Onion: Peeling Open Vocab Object Detectors Under Noise cites this paper.

Robust Onion: Peeling Open Vocab Object Detectors Under Noise CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:15:50.059824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T00:42:36.605869Z digest=sha256:c6302f8624d3f3d0771bb81165764bb85945c7e6324ef77beae359da14bde153

Observation 4053493b-0866-4e49-b640-edcdf83ca423 · inbound

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions cites this paper.

A Reconstruction-Based Framework for Caption Evaluation Beyond Reference Captions CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T00:03:54.834315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T00:03:54.834315Z digest=sha256:2f8cae8bb5851a02983edae13353e2ac02058f2cc8227d97844e08a4d392b3fa

Observation e6977576-efe3-4033-b901-1ae2ef49a408 · inbound

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding cites this paper.

Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding CapArena: Benchmarking and Analyzing Detailed Image Captioning in the LLM Era

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T23:42:35.033889Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T23:42:35.033889Z digest=sha256:638e6ac69af84979a9790c339c1fdb77bd251c4f186f9f49d2f850ef498d5dac