Pith. sign in

Paper Citation Record · LEDGER

Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2403.09333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.09333 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T11:19:33.923271Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T12:13:16.249272Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ce95c1e9-0208-4384-a744-d0282c005580 · inbound

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding cites this paper.

ChatRex: Taming Multimodal LLM for Joint Perception and Understanding Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-12T11:19:33.923271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:19:33.923271Z digest=sha256:c08dade00e794ce4526cea938a3a62fb97aa5aa74e2852f7dca46c4c4fc27394

Observation 73b3437d-0305-4477-9b43-6b4fcb899615 · inbound

Apollo: An Exploration of Video Understanding in Large Multimodal Models cites this paper.

Apollo: An Exploration of Video Understanding in Large Multimodal Models Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T16:11:10.743463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:11:10.743463Z digest=sha256:d0a133201d595567d014684f6defda9ff222742cfb3260b5753b4618b725a4d0

Observation 072ecc17-1522-4120-aff8-b9e5849f967c · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.337047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.337047Z digest=sha256:52d538e46bb3f80a22a99f525c73d147e286329448c6823e68853dff360ddabc

Observation 407a7a7f-e09c-4a35-870f-65152987203c · inbound

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models cites this paper.

Migician: Revealing the Magic of Free-Form Multi-Image Grounding in Multimodal Large Language Models Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-10T21:10:40.708082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:10:40.708082Z digest=sha256:58f0040a59d4df5af0aad04f4fbd3b587c73a827859eb18a62044ec10bff118a

Observation 57b661ce-687e-4bac-9ddc-1c6259d63f24 · inbound

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning cites this paper.

UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:56.613151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:56.613151Z digest=sha256:f5c1252ef380f3c28ba0ed11cd3d257e2658ef14c44b51e1280a30bbb0a4d833

Observation 0bcc7138-aea4-43ed-945d-f88edb4b5338 · inbound

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning cites this paper.

Rex-Thinker: Grounded Object Referring via Chain-of-Thought Reasoning Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T10:55:14.895256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:55:14.895256Z digest=sha256:ebc7145c3b4b21633ab33d5e9ea9c70efac25a59c50b26d680f869430d897e18

Observation 765e1542-dedc-463f-bd2e-ee305f9bfd26 · inbound

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning cites this paper.

Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-19T06:52:07.957116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-19T06:50:02.607136Z digest=sha256:cb0c8a9c97a53ea385a850010b866e39da1425a87bd3080ae07d90152220dbaf

Observation 9d2ed155-c7a8-4ddc-ab93-2b9ce92be1c6 · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding Griffon v2: Advancing Multimodal Perception with High-Resolution Scaling and Visual-Language Co-Referring

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.251440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:687bb75cfdc1eddf760290565ee48b234563924f73c28addb740a617a8f40a96