Pith. sign in

Paper Citation Record · LEDGER

Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

As of 13 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2406.16866.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.16866 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:31:37.035027Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T03:06:30.418964Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c015d103-93b1-4f2f-8f8d-eeb9445e64a7 · inbound

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs cites this paper.

MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T14:31:37.035027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:31:37.035027Z digest=sha256:2e2573e3088b07330a4242b437898a78bece2c19ad9f2756a4f4cde4924d3319

Observation 8847315a-6971-4b00-a272-49be1b9d1b92 · inbound

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models cites this paper.

CaLoRAify: Calorie Estimation with Visual-Text Pairing and LoRA-Driven Visual Language Models Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:36:34.710222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T16:36:34.710222Z digest=sha256:52ee1be9cb8f83eabdbfca5c9872a4897c6f516aaccab803068b9b765ceed876

Observation 0b2efd09-6a65-4065-bb2f-bb74d5075251 · inbound

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer cites this paper.

LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:59.433530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:59.433530Z digest=sha256:f1e91f2f855e72ace0157afdddfcc5a5f7e8191cf3713ffc68157090cf5b0f83

Observation 05ff9b75-378d-4771-9df1-bd40edd1ee85 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.476383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:54df51a3f3fbd525b65aa51e53da6beef449430fd3c9126961905363fe13a032

Observation b62095e3-85b8-401e-96ee-90db625f3260 · inbound

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training cites this paper.

LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:53:26.615927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T10:53:26.574350Z digest=sha256:3462bab06703b5f1a4b07cda2f6feb6838a7cb9dee0966cce301ab900a82d56f

Observation 0d56b3be-ae1c-4d34-9335-5e815e5e7573 · inbound

The Mechanistic Emergence of Symbol Grounding in Language Models cites this paper.

The Mechanistic Emergence of Symbol Grounding in Language Models Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T09:45:01.649392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:45:01.649392Z digest=sha256:bfa9325d39ac95b59e99830809b17438b8b12e0dc18a0a90366fc3c139c91687

Observation 7f446895-3ac8-403a-9351-5e3c71dd2064 · inbound

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs cites this paper.

FindIt: A Format-Informed Visual Detection Benchmark for Generalist Multimodal LLMs Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:30.421473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-28T10:16:53.915716Z digest=sha256:7c5ebc9e1108d6c3d7dbcecad24c68f7e4ba54f0e15558e5df8544ca14fa1803

Observation a2cf8f59-6714-45bc-8395-4bdbb072b083 · inbound

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding cites this paper.

HKVLM: Faithful Query--Region Binding for Frozen-Detector Visual Grounding Revisiting Referring Expression Comprehension Evaluation in the Era of Large Multimodal Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T11:14:52.015620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:14:52.015620Z digest=sha256:77b08b65fe8d8fc8c39fb451d498af0b859f25c56da66d35cc90dbc28a74ca37