Pith. sign in

Paper Citation Record · LEDGER

VLRM: Vision-Language Models act as Reward Models for Image Captioning

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2404.01911.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.01911 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:15:04.420614Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:12:34.682493Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 811f7398-0b2c-445b-b4d2-9ad1df5ff3ae · inbound

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models cites this paper.

ZIUM: Zero-Shot Intent-Aware Adversarial Attack on Unlearned Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T12:15:04.420614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:15:04.420614Z digest=sha256:9a47bec5a12f11ac1a67b992183c73ae00cc673ea51f4d5b88d9afda677cafbe

Observation 25197f98-ce85-42bd-b7fc-34da4153e12a · inbound

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models cites this paper.

WSVD: Weighted Low-Rank Approximation for Fast and Efficient Execution of Low-Precision Vision-Language Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T21:08:18.007439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T21:05:09.254086Z digest=sha256:53549fc97b7da8ed58cf0e03194299dbe2c6d96aff1b03609b749126326d5dbe

Observation 51cd10fd-0a78-4d54-9f02-5d980671c7ef · inbound

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers cites this paper.

Multimodal Backdoor Attack on VLMs for Autonomous Driving via Graffiti and Cross-Lingual Triggers VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.573580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T20:30:29.032353Z digest=sha256:eb8c6a117e5687208126746013b5e11cd9e2a1bc10761acdcfa1d7ea40145809

Observation c02c1e64-9071-4be2-b368-8855798c18e6 · inbound

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation cites this paper.

DREAM-S: Speculative Decoding with Searchable Drafting and Target-Aware Refinement for Multimodal Generation VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:02:34.569828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-28T18:55:51.474956Z digest=sha256:f9db4e3a75bcad9302089331f5e7ad13a5e691339d511cdf816095541bfe3c8e

Observation 08a734bf-3648-48ef-8de3-8defe27e8707 · inbound

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models cites this paper.

LASER: Loss-Aware Singular-value Decomposition and Rank Allocation for Efficient Low-Precision Vision-Language Models VLRM: Vision-Language Models act as Reward Models for Image Captioning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:12:34.684066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T19:09:29.347284Z digest=sha256:5fa1a79491db0aaccfdff4721a1e3d09b4ab4b626310435a06567b87baa8e8ea