Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 24 inbound Pith citation observations for arXiv:2404.07973.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:55:14.903410Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T17:18:43.769941Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 52c46c55-579c-46fe-97ec-d0cf92d06fcc · inbound
PaliGemma: A versatile 3B VLM for transfer Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 163
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bf729760-af01-45a6-893f-d69d7bddf096 · inbound
Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 298
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb2197b6-7ad7-43e4-bd0f-417a2b0b5481 · inbound
DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 109
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d4e3094-5d35-4e77-99a9-2c7866e3cecf · inbound
Qwen2.5-VL Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f3329cb4-26ea-4ff8-b71b-a584ec2106a5 · inbound
InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19323831-a40a-49de-b004-6b7c50a0a822 · inbound
Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae642888-c3de-4cbe-aa24-3161505b0301 · inbound
Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 061cff8b-33b5-4e83-a9cd-f71616a5ae0e · inbound
SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf03d328-6614-4812-9942-6988fa4739e2 · inbound
Mitigating Object Hallucination via Robust Local Perception Search Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e50f82b-d859-4e4a-b256-471307846976 · inbound
Region-Level Context-Aware Multimodal Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c4736ca-60ba-4593-9f5a-14e5318b593a · inbound
InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 174
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 65a71161-201a-4948-826e-c30a6127c282 · inbound
Strefer: Empowering Video LLMs with Space-Time Referring and Reasoning via Synthetic Instruction Data Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34366149-934d-4980-ac63-399f83fe1d0c · inbound
Grounding Everything in Tokens for Multimodal Large Language Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e0166ae-07b0-46e7-b378-14b89929b444 · inbound
LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 224
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f6556e68-116d-4302-a7a0-ae0bb9f0cddf · inbound
APRVOS: 1st Place Winner of 5th PVUW MeViS-Audio Track Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 79c338b0-a657-400a-aeeb-d4f246691794 · inbound
AgentRVOS for MeViS-Text Track of 5th PVUW Challenge: 3rd Method Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7edf8782-c164-4990-aea0-70f39a55e4cd · inbound
SpatialForge: Bootstrapping 3D-Aware Spatial Reasoning from Open-World 2D Images Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3c192de7-c3cf-4579-bbb6-8210ba4bad56 · inbound
SceneParser: Hierarchical Scene Parsing for Visual Semantics Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51c62cd9-08e9-43a4-8844-3cd356ccf0fa · inbound
See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19cf84d4-0dc3-42cc-891a-1538e68afda9 · inbound
Qwen-RobotWorld Technical Report: Unifying Embodied World Modeling through Language-Conditioned Video Generation Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 50f680ef-28ec-4ccd-8347-75f771161114 · inbound
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8378c805-3cb9-4fde-afd2-956cd2c9cfa9 · inbound
Qwen-Audio-VAE Technical Report Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 194
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81988803-75fe-4b14-a681-bdc1f7221ccd · inbound
CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577a0dee-2ba4-4ba8-b0ae-991bf551e466 · inbound
Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.