Pith. sign in

Paper Citation Record · LEDGER

OmniCaptioner: One Captioner to Rule Them All

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2504.07089.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07089 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:58:36.852405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:17:51.543739Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e91735c6-df67-4819-b970-c7a0857b77f5 · inbound

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning cites this paper.

SC-Captioner: Improving Image Captioning with Self-Correction by Reinforcement Learning OmniCaptioner: One Captioner to Rule Them All

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:58:36.852405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:58:36.852405Z digest=sha256:a0c88cf4ae8f5c4e549eab04e279d45b8c731d23704aacbaf212d1272ef449db

Observation a7ed7b5c-c64b-4b29-89e0-41dc41e350d9 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards OmniCaptioner: One Captioner to Rule Them All

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:23.473947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:23.473947Z digest=sha256:c5f2c259a049ab20eeed51dbd6741681cd8cd6e79fecec358a3298b51b762797

Observation c51d8ef9-8ad0-4d17-9ee5-a3d61fc27f41 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniCaptioner: One Captioner to Rule Them All

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:08:37.265735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:4e398a627f69d701a3a8040e914013fd228e6263f790777294072282d29d98cc

Observation c21bfeb1-2a69-48f2-8e4d-de3b6d031cfa · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniCaptioner: One Captioner to Rule Them All

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.786122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.786122Z digest=sha256:645fe40dd65743c00804143727275b84f124fb1f590bf1b8e5b7332d3752861a

Observation 661e7ce4-ad9a-45d9-bd79-4b1ae5053281 · inbound

CodePercept: Code-Grounded Visual STEM Perception for MLLMs cites this paper.

CodePercept: Code-Grounded Visual STEM Perception for MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T23:22:13.847876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:22:13.847876Z digest=sha256:15da5db1f5705c8bde346b77874b061c875d6cc01274470e688b0af9036dda72

Observation 30f33aa0-b206-4d6c-93c1-2d1309a7affe · inbound

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions cites this paper.

DetailVerifyBench: A Benchmark for Dense Hallucination Localization in Long Image Captions OmniCaptioner: One Captioner to Rule Them All

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T00:10:52.847901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T18:39:29.592129Z digest=sha256:6cb89cf50666333dc2f1b718205bb11d4d26265d632c83b172e18017fa134454

Observation 7fbaec1f-7bdb-4f93-a552-67b33733d204 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning OmniCaptioner: One Captioner to Rule Them All

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:55:44.238614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:ddcae9c438385632c77dcbdd35322d0973de4eba994be4fac9016053cd81b577

Observation f4e8a51d-2404-4008-af5e-42bb65278d08 · inbound

Watch, Remember, Reason: Human-View Video Understanding with MLLMs cites this paper.

Watch, Remember, Reason: Human-View Video Understanding with MLLMs OmniCaptioner: One Captioner to Rule Them All

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:27:15.021888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-27T22:00:28.350003Z digest=sha256:37460940dec6c446a6eafb85ec6120d64251d747a9b290f329d8e0b0ecebd5c5

Observation 2f302196-1973-4509-88fc-adeaf37cf13d · inbound

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models cites this paper.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:17:51.572452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T03:11:49.713677Z digest=sha256:3b2f1b879cbfac15dd63da9d7813dee1b5de7b5aa90c9de4f6bed03116be64d9

Observation 28cde4be-7dcf-4d3d-893c-76029b9f4b60 · inbound

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models cites this paper.

Scene Graph Thinking: Reinforcing Structured Visual Reasoning for Multimodal Large Language Models OmniCaptioner: One Captioner to Rule Them All

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T16:12:58.882161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T16:12:58.882161Z digest=sha256:2c9303dfc5475d8a57de37a633ab03a1ca43ff947ca6fcc9ca0b4136558fc2a2