Pith. sign in

Paper Citation Record · LEDGER

ChatterBox: Multi-round Multimodal Referring and Grounding

As of 23 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2401.13307.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.13307 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:52:30.302929Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.009759Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation b459bffb-b2d8-49b4-b2f8-581389ed9697 · inbound

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM cites this paper.

VideoRefer Suite: Advancing Spatial-Temporal Object Understanding with Video LLM ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T22:52:30.302929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:52:30.302929Z digest=sha256:0e6011a596b883d768a64a6c3ef0a1867e218a0e2d89b77299afd05b58784489

Observation 0416a874-b557-45af-b665-202975af1755 · inbound

Visual Large Language Models for Generalized and Specialized Applications cites this paper.

Visual Large Language Models for Generalized and Specialized Applications ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-10T22:08:09.272318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:08:09.272318Z digest=sha256:0de9648753c6bca5c1e66aa69a7b7aa5a4195050dae8975565ed89b3d2a36e0c

Observation 5f725db5-be10-455e-9c8c-8492c98d1363 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.891664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:2a9fb327070d8b1bd02aebd7423be1f4804ed39a5ecd2eac589da82f3917eaea

Observation 01f06659-eee1-4f52-a87b-5cf22220bed6 · inbound

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding cites this paper.

LongVideo-R1: Smart Navigation for Low-cost Long Video Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:01:33.380638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-15T20:01:31.129959Z digest=sha256:69885a637c9573e94070d21fa18528183ea841d906b62272f4edd9842fe32af2

Observation 665d09d2-0384-475b-81e5-6c6e3077e78b · inbound

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding cites this paper.

See What I Mean: Aligning Vision and Language Representations for Video Fine-grained Object Understanding ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:13:16.194382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-05-20T12:10:54.874012Z digest=sha256:5391f79456d372bbefdc3b1bde5199f4f8183912b060bfcf55105e952d2d6d85

Observation b6219926-56a0-4a5a-95f8-3b4ec1a0777a · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ChatterBox: Multi-round Multimodal Referring and Grounding

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.011788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:adb5d6e11ef3d71cfe22e421518800ba43bd138b19f3e53e21b7c81ecb381ac5