Pith. sign in

Paper Citation Record · LEDGER

GRiT: A Generative Region-to-text Transformer for Object Understanding

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2212.00280.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2212.00280 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:50:59.184754Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:26:26.561841Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 52ccaa0b-309f-4256-bf89-57f0f0adba60 · inbound

VideoChat: Chat-Centric Video Understanding cites this paper.

VideoChat: Chat-Centric Video Understanding GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:30:00.669366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-13T23:30:00.457974Z digest=sha256:ed333b1f614a6fa46ed52f35fd18584d7eb4d97535fc1ea9bc6d57f8c032f0c6

Observation 9f63ba52-0c1a-4cd7-8685-50d70cb3c826 · inbound

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension cites this paper.

SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T16:59:50.569621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T16:59:50.495335Z digest=sha256:acc47273ddecfb316ca278bc798fa9fd02a286c80dc24cbe3116d2cb26c52a1c

Observation 5de9c206-1048-4b36-aada-b53e3f9bf0e4 · inbound

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) cites this paper.

The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision) GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 138

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:26:06.441967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T23:26:06.183574Z digest=sha256:2b1d4e3819c8d4181ca6ce95abb38733d2f0d6e59baabe7d00da2079b2ecb76e

Observation b31bb02d-b0fa-4232-b72a-972c99825928 · inbound

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models cites this paper.

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:41:30.576293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-07T04:13:26.433133Z digest=sha256:d1b38c0cb7d37c80e1fdba824e1821481a537ae8e530b09cbad29839c4598677

Observation 2f15b9b8-83ff-41a7-baae-17fe97c41c8d · inbound

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models cites this paper.

WorldJen: An End-to-End Multi-Dimensional Benchmark for Generative Video Models GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:00:37.412332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-08T19:08:07.487384Z digest=sha256:7923b1f58d4e00da53d52b4a25e45b3c37254ae892ba7745a02ba51be9b08680

Observation 151fbe20-44ac-49a7-bd18-ca287f536b89 · inbound

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models cites this paper.

Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:26.563750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T10:57:09.699207Z digest=sha256:6bbc270b2a5a3fee76595566cbba30bd464b114dfbf6a1fade75c2f1b2edc26a

Observation c61e9bcf-d9d8-4974-b9e9-18f4fb24598d · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World GRiT: A Generative Region-to-text Transformer for Object Understanding

Reference 157

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:b792dc5b8b206bb433c829b8aac552a3da55a7637e304d07d2a213d6199f5a2b