Pith. sign in

Paper Citation Record · LEDGER

Generative Multimodal Models are In-Context Learners

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 17 inbound Pith citation observations for arXiv:2312.13286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.13286 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T15:24:46.417390Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ae3d4092-7407-449d-8aa4-9e8916771767 · inbound

CogVLM: Visual Expert for Pretrained Language Models cites this paper.

CogVLM: Visual Expert for Pretrained Language Models Generative Multimodal Models are In-Context Learners

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:46:06.530109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T15:46:06.334088Z digest=sha256:e476208407971860c414403f302f850dd16afa529ddd80a88eb09e8d11e32deb

Observation 6825f400-1343-4766-84a6-68c9cac61b92 · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI Generative Multimodal Models are In-Context Learners

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.625490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:34ed9ae8ee3d0cda43b0c0bd546ef0b6334a3de68716f972217cbf1da9aa21f1

Observation b0911ed6-c864-4058-885a-60fdd5a49a52 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Generative Multimodal Models are In-Context Learners

Reference 105

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.258789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:d65742473fca73db3db231403972200919c6e28d7eb9813ff1f9dec10a02abf2

Observation a979d1ca-2f30-44fb-abeb-5bd5a1198235 · inbound

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models cites this paper.

Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models Generative Multimodal Models are In-Context Learners

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:44:47.444102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T07:44:47.355960Z digest=sha256:d13d2b578338d802a9f28bdc73794526eb984da3aa7614c6b2b211e1f1bba237

Observation c634f635-b37b-4346-92e7-85eb92dcf127 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Generative Multimodal Models are In-Context Learners

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.198403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:14778f15ccca389caa408782b9d9789f51a9349970354d567b02a54045d46ca2

Observation 9a5d5d9d-7d52-4730-b776-94eaa623f032 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Generative Multimodal Models are In-Context Learners

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:09:30.412895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:2bc021c410c3f37fc1fd1d7ded297f92c63ad097e06a74c116ba266cc45e456f

Observation 61279890-74ee-4548-bfe9-d7e7817849f0 · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Generative Multimodal Models are In-Context Learners

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.569913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:583b76e7ad8e9533b24a106036131e86e914f4cb70c0135073d9802a8aa06e75

Observation e6da27c2-5c8e-4e1b-a7ba-711c96c4e96b · inbound

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation cites this paper.

Show-o: One Single Transformer to Unify Multimodal Understanding and Generation Generative Multimodal Models are In-Context Learners

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:03:33.563332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T21:03:33.427939Z digest=sha256:57ce76695697dd62bef3900bac8fc3431128b9e0867f33bbb2c4539297415e31

Observation 110289f0-5ca2-4927-8b8a-e23636be0980 · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths Generative Multimodal Models are In-Context Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.417390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.417390Z digest=sha256:c912432a39d34662204b39d81c2633ea251c253fd550e2a342d0c7299e7c49da

Observation 93fee7ad-4953-479f-a040-2b380c6b26f9 · inbound

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models cites this paper.

EVEv2: Improved Baselines for Encoder-Free Vision-Language Models Generative Multimodal Models are In-Context Learners

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-08T14:25:55.834228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:25:55.834228Z digest=sha256:01ed9b81f731727b88b829cc6b9f95436592c964fb3e1e8ef4a49715970c1d56

Observation ec11eb0b-6777-46b3-b61b-780cc315fee6 · inbound

MMaDA: Multimodal Large Diffusion Language Models cites this paper.

MMaDA: Multimodal Large Diffusion Language Models Generative Multimodal Models are In-Context Learners

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:50:59.902495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T14:50:59.661153Z digest=sha256:34e9f5aa37f5a20b28b089280d00c7aeada58749a5428b85d338563e4479161f

Observation c2c0e2d7-3117-4f63-9400-0e3e1d0fe880 · inbound

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation cites this paper.

R2I-Bench: Benchmarking Reasoning-Driven Text-to-Image Generation Generative Multimodal Models are In-Context Learners

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:48:55.005689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:48:55.005689Z digest=sha256:d53e0897cba28e1a005ac45e447f2e4d9d833b421ced8b0709c1a880488ef899

Observation 6e9ff967-d6ae-49d2-9586-7c1815053cbd · inbound

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools cites this paper.

On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools Generative Multimodal Models are In-Context Learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:27:36.829877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:27:36.829877Z digest=sha256:ba54160f65f28f5f32a0921b6c2bf5047b6768707c9fdb7fd8004ff72c8aeca6

Observation 4da437f8-b144-4448-a08b-69ee1b9052b8 · inbound

Grounding Everything in Tokens for Multimodal Large Language Models cites this paper.

Grounding Everything in Tokens for Multimodal Large Language Models Generative Multimodal Models are In-Context Learners

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-16T23:31:21.916857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T23:31:05.422935Z digest=sha256:b39879b2721187ff892244338d487b8032295ae2a725e40a34e597fea031a490

Observation 7e283520-1b01-40ee-84a0-82bf6d97e28f · inbound

ReMoT: Reinforcement Learning with Motion Contrast Triplets cites this paper.

ReMoT: Reinforcement Learning with Motion Contrast Triplets Generative Multimodal Models are In-Context Learners

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-02T19:58:25.979720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:58:25.979720Z digest=sha256:8ab4b4cf6e30da15deba97243ed9d191f49b851bb72442b5e2cf484c2051213c

Observation 756d0efe-bc9d-4e89-99d9-89e0a93114cc · inbound

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning cites this paper.

HPP: Hierarchical Programmatic Probing for Long Video Understanding by Decoupling Perception and Reasoning Generative Multimodal Models are In-Context Learners

Reference 245

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T06:39:37.521811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T14:19:53.450263Z digest=sha256:5c8bd18fb1f68f6ea55453eaeaadaceaf2f27f88fe3f971f2192d55805c03f8c

Observation d96f5082-d825-49ea-9723-2d8cf0e31c10 · inbound

Qwen-Audio-VAE Technical Report cites this paper.

Qwen-Audio-VAE Technical Report Generative Multimodal Models are In-Context Learners

Reference 154

Resolution
unresolved
no resolver link, observed 2026-07-14T03:31:19.309532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T03:31:19.309532Z digest=sha256:c335ad2ad4bc2f2f8a630cd28f59a325bd4ded56446069fea41135c14bc58d72