Pith. sign in

Paper Citation Record · LEDGER

Fine-grained Image Captioning with CLIP Reward

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2205.13115.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2205.13115 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:54:07.333485Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:41:30.171123Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 82ae280a-d73c-41e1-9c54-c9f95123f773 · inbound

The Double-Ellipsoid Geometry of CLIP cites this paper.

The Double-Ellipsoid Geometry of CLIP Fine-grained Image Captioning with CLIP Reward

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T15:27:08.719025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T15:27:08.719025Z digest=sha256:fc109dc1611dd97db99356e4b141a512c895824a32bd0667734853841e453ac7

Observation 983bf124-daaa-4ea4-b362-a13ce4860e92 · inbound

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding cites this paper.

Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding Fine-grained Image Captioning with CLIP Reward

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:03:20.260110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:03:20.260110Z digest=sha256:bdff8f244b5035df96ba898d7597f7318847ae504609d5836e097e2df518ac8e

Observation 276f4202-75b9-43d5-b962-9ca78f2917a6 · inbound

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning cites this paper.

InverTune: Removing Backdoors from Multimodal Contrastive Learning Models via Trigger Inversion and Activation Tuning Fine-grained Image Captioning with CLIP Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:58:06.678094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:58:06.678094Z digest=sha256:7ec37a1b3fe0a3b47d5ab811efdf9f26b730d21f0f2062c5c04828431f15bcb8

Observation ce59b7fb-8c53-4bfc-b2f0-6ba1f2fa19ba · inbound

The Devil is in the EOS: Sequence Training for Detailed Image Captioning cites this paper.

The Devil is in the EOS: Sequence Training for Detailed Image Captioning Fine-grained Image Captioning with CLIP Reward

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:54:07.333485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:54:07.333485Z digest=sha256:bb50d70bd81373f78961c7d597e82aaede0648ef96489add002e65012577c030

Observation fae5e868-8fd5-43e6-92a9-ed7f29afb72c · inbound

Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens cites this paper.

Modality Bias in LVLMs: Analyzing and Mitigating Object Hallucination via Attention Lens Fine-grained Image Captioning with CLIP Reward

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T05:04:22.275679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:04:22.275679Z digest=sha256:5f7f501f7fba335a1f01061d3864af61ff69ae1367ff65ccc4aa71563be6156f

Observation 9fd07743-f4d3-4744-b35b-0ba5248616c4 · inbound

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs cites this paper.

Long Story Short: Disentangling Compositionality and Long-Caption Understanding in Contrastive VLMs Fine-grained Image Captioning with CLIP Reward

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:41:30.173527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-18T14:41:20.403259Z digest=sha256:7716a940d3679a1dac43087b99c3cd6f07edba6eae4bb5189d5e7b3dbaa942a1