Pith. sign in

Paper Citation Record · LEDGER

Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2406.16866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.16866 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:59.433530Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:30.418964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 0b2efd09-6a65-4065-bb2f-bb74d5075251 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.433530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.433530Z digest=sha256:ff269e170eada7cab556f62e2a16200cb78abe96b87c2933ead8ac0918f2ad9e

Observation 05ff9b75-378d-4771-9df1-bd40edd1ee85 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.476383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:d3ecd6111e049f10c398da93e38bfe51774d1c75c35ad9acaafe1638fb1966c5

Observation b62095e3-85b8-401e-96ee-90db625f3260 · inbound

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training cites this paper.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.615927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:1d09c223902d6d2740ae6ad2b8c3cf7938ad7482b0906a3f5ab7d142b90cf410

Observation 0d56b3be-ae1c-4d34-9335-5e815e5e7573 · inbound

The Mechanistic Emergence of Symbol Grounding in Language Models cites this paper.

The Mechanistic Emergence of Symbol Grounding in Language Models Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:45:01.649392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:45:01.649392Z digest=sha256:05ee1a8ef0f1d0b1880a4e4c792926d614787a954c51eaf8ae2f963c7551d9eb

Observation 7f446895-3ac8-403a-9351-5e3c71dd2064 · inbound

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs cites this paper.

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:30.421473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:16:53.915716Z digest=sha256:5f866e6e180ddafd8a84ec2c771b3787cc49c30e2691bbb3b10e670dcd37007b

Observation a2cf8f59-6714-45bc-8395-4bdbb072b083 · inbound

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding cites this paper.

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:14:52.015620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:14:52.015620Z digest=sha256:c0506aa714538dd96d093f4fcd247426146c3ad0d8e377277a5482754eea8042