Pith. sign in

Paper Citation Record · LEDGER

Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2404.07973.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.07973 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 24 of 24 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:14.903410Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:18:43.769941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52c46c55-579c-46fe-97ec-d0cf92d06fcc · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 163

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:21.316983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:3ba16a7f1618088fbf3215a8b070fa7cd9e015e9375b52f65cab301ca68f4b71

Observation bf729760-af01-45a6-893f-d69d7bddf096 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 298

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.322854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:8a5826e30d5c0b5f0edb4763f1503b024511297659285fef11fe7de8c3f86b2a

Observation eb2197b6-7ad7-43e4-bd0f-417a2b0b5481 · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:26.618781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:109c2651958e5f2013b02faeddd79e7c9efc41989fff2f61ae8607bfd8fe3f51

Observation 6d4e3094-5d35-4e77-99a9-2c7866e3cecf · inbound

Qwen2.5-VL Technical Report cites this paper.

Qwen2.5-VL Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-23T02:25:19.039848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-23T02:25:04.405036Z digest=sha256:707906d0df8dd48355bc42f7c204b583d855bde41d446e746f657816a4fc6e97

Observation f3329cb4-26ea-4ff8-b71b-a584ec2106a5 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 145

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:41:08.173791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:95f970afdae7187ac7ee492fe542885eb4fe93857912ffcb05a9b40b5093d2d5

Observation 19323831-a40a-49de-b004-6b7c50a0a822 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.903410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.903410Z digest=sha256:f03a1fea6227fd70897332dda4701f897097548b0fc4059d2eeb046c834328e5

Observation ae642888-c3de-4cbe-aa24-3161505b0301 · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:14.279027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:14.279027Z digest=sha256:b1483f25f8a9f2e4b4f08dcf47b9b2b55ea587c25f3f5a5ee21a4ef13de05b74

Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · inbound

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing cites this paper.

SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:50:52.768322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:50:52.768322Z digest=sha256:3f8e8c90e80d2087314d140382223919576eed67f6717b2e312597f82caee3c3

Observation cf03d328-6614-4812-9942-6988fa4739e2 · inbound

Mitigating Object Hallucination via Robust Local Perception Search cites this paper.

Mitigating Object Hallucination via Robust Local Perception Search Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:54:03.713136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:54:03.713136Z digest=sha256:1aca2cf5c48c6b85b12d7ffca7acb7c690c177bb078c4d6417e0b9e1dcf0b13b

Observation 4e50f82b-d859-4e4a-b256-471307846976 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.130365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.130365Z digest=sha256:035b515e38d261c9486d526b81e5ba3ae22e79cb9ec077b5adfe9bff9fc2d1e8

Observation 7c4736ca-60ba-4593-9f5a-14e5318b593a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 174

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:58:58.946719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:89a16c3e46afcfc13e24836d8197f7af49ce092d3390ea02145754d63bfd912b

Observation 65a71161-201a-4948-826e-c30a6127c282 · inbound

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data cites this paper.

Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T10:58:19.257522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:58:19.257522Z digest=sha256:508eaa0f958cc0ef7dc772b6e6d9f6b9821e35d6dfc7d4df5ed4d4803c0281e3

Observation 34366149-934d-4980-ac63-399f83fe1d0c · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.900012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:d001ddd515574683b6961376b7249833eff6eb856b68efb1fd4a4f64798d7e97

Observation 8e0166ae-07b0-46e7-b378-14b89929b444 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 224

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:08.196499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:4ffca3e658b4181a3ec6a0d13a16c2f99340289a7249b56c67f41762fbaabcd1

Observation f6556e68-116d-4302-a7a0-ae0bb9f0cddf · inbound

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track cites this paper.

APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T03:29:21.572846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T03:29:12.783993Z digest=sha256:d7354a54262fa6127000d373cb708fa081a3d6110e776fc27676103120c8e970

Observation 79c338b0-a657-400a-aeeb-d4f246691794 · inbound

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method cites this paper.

AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:33:41.834379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T05:14:25.423302Z digest=sha256:d9fd24cd214b9b3e333dab8233746497f69077d04fb0e6c0331d5fc1a43600dc

Observation 7edf8782-c164-4990-aea0-70f39a55e4cd · inbound

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images cites this paper.

SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:27:01.649239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T01:26:47.052025Z digest=sha256:c433aefe22d7ce335a8a682a4427c7b79a982c5e6dd88026172f5c6f36a6c7df

Observation 3c192de7-c3cf-4579-bbb6-8210ba4bad56 · inbound

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding cites this paper.

SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T20:55:04.145207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T20:51:56.131205Z digest=sha256:a8523f6bea1ce77b38736ada93f3557b3ef9ffb692a4b2169fae09b8b0427d93

Observation 51c62cd9-08e9-43a4-8844-3cd356ccf0fa · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.197690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:275c6f790e7482bc7ade0829e43d94b443eb617664480b1a26e7c37e365a0c90

Observation 19cf84d4-0dc3-42cc-891a-1538e68afda9 · inbound

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation cites this paper.

Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 181

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:18:43.771463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T04:19:26.332718Z digest=sha256:70ba1dec69dc94bbfa389463465ae750e0de03724115167eb582b90a5a0e2598

Observation 50f680ef-28ec-4ccd-8347-75f771161114 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:972e8380cab42b087d08ca6f21e65bca9a47cc8ccb161b05ad8e1073cbade893

Observation 8378c805-3cb9-4fde-afd2-956cd2c9cfa9 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 194

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:1cc7f676acb3157a1bd47ad076d92ec51a297c7480d1e02657c7f112875e547d

Observation 81988803-75fe-4b14-a681-bdc1f7221ccd · inbound

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement cites this paper.

CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-05T10:23:58.883843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:23:58.883843Z digest=sha256:ba81934addee59134b66b1f9d2f24065cb5580e906f44362b0fd9bd43f4f0506

Observation 577a0dee-2ba4-4ba8-b0ae-991bf551e466 · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:45.544126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:45.544126Z digest=sha256:0b2d252b7038898496b551d197326cc4695ab64ed58bf93c09fc93e87f450ec9