Pith. sign in

Paper Citation Record · LEDGER

Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 26 inbound Pith citation observations for arXiv:2309.02591.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.02591 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:50:00.168507Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

27
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 739d5dad-b6b6-47d5-99ac-9a6729b8ee4d · inbound

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation cites this paper.

Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 214

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T20:06:44.615691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T20:06:44.480769Z digest=sha256:9791a9000884b94cc40708daf1e045a5a649aec375b03f6d083550cabea01667

Observation 2e63b695-acb3-402b-8270-6b1634298980 · inbound

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation cites this paper.

SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.344214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T22:48:36.010306Z digest=sha256:003ea5eaea833c0e13ba0630d9906fb65cbf06171439c6fec3d96017621ecee4

Observation a157ff30-ed01-4622-8d76-af1ffb73c384 · inbound

Chameleon: Mixed-Modal Early-Fusion Foundation Models cites this paper.

Chameleon: Mixed-Modal Early-Fusion Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:03:28.152026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:03:27.919346Z digest=sha256:93262908c10261a08553ba0d5bd184c67cc73af365eaf58898397dbe6f21c5a8

Observation d85a6cb4-c31f-454c-a2d5-09905940a6cd · inbound

PaliGemma: A versatile 3B VLM for transfer cites this paper.

PaliGemma: A versatile 3B VLM for transfer Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:10:20.609481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T13:10:19.972353Z digest=sha256:2abf9758a39ab13735fb3eccd0005be969a738b56e02b2f187ba7a3fc7f44190

Observation 3cdb2fec-7473-4b18-a944-37d7bf23c709 · inbound

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model cites this paper.

Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:26.997276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-13T05:57:26.887069Z digest=sha256:64c1c373c32f19b3f8ab01ad334fe43ff22edbc4a321abcf9fcbdbf7ba261dbc

Observation 50606391-af01-487d-a526-855b56149c58 · inbound

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation cites this paper.

VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.382770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T00:26:21.313005Z digest=sha256:f525328c82438fba729788e85fc7057101474f4ab6cceaf89038354c962ec311

Observation 7554d3b9-a55c-47e6-babf-3ee709a2ed4a · inbound

Emu3: Next-Token Prediction is All You Need cites this paper.

Emu3: Next-Token Prediction is All You Need Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:09.455699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T10:56:06.418360Z digest=sha256:5d549b8da1bb4d17c13ba335f5e7628bf87e3aa17ea09227b1cf834a9b6b30f2

Observation 5d99ca5f-b621-4cb0-8da2-388754be6df9 · inbound

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models cites this paper.

Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T02:48:45.089761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-18T02:48:44.900467Z digest=sha256:5a776d40c317a2e74dab0c2b60ee59f56700d2569b70a424b4b5d43645f65dff

Observation a911f41a-9533-4efb-ade8-38a6107d7050 · inbound

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies cites this paper.

DualToken: Towards Unifying Visual Understanding and Generation with Dual Visual Vocabularies Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:52:16.729035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-22T23:51:43.934329Z digest=sha256:49edae1377d57698535098edda4f941b585220fef4ffaf032d570186bc73a7ac

Observation 1f6f407e-d188-444d-b36f-0b6e4a9b9b6d · inbound

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation cites this paper.

Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:04.730499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T07:24:04.460276Z digest=sha256:e22445851477e08eea4f70f400d4aca86e6f338dfd56c7adcb9efc25668029f0

Observation 0349970a-a02d-4bfb-b826-ce4611bc89f9 · inbound

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation cites this paper.

AR-RAG: Autoregressive Retrieval Augmentation for Image Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:50:00.168507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:50:00.168507Z digest=sha256:497f2c9c822b73b06db75886b0e8b748e6450ff31431ea34e80f3fabbb457ea5

Observation cddd270a-1f4d-403d-a095-4b938064210e · inbound

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems cites this paper.

CuRe: Cultural Gaps in the Long Tail of Text-to-Image Systems Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:27:16.282820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:27:16.282820Z digest=sha256:05fd8512a0154f5cfc7face2af46e343bac9e16107ce3b6e5669840ea452ed30

Observation e9b4edd9-af5c-4184-8228-db135bf16a19 · inbound

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation cites this paper.

Pisces: An Auto-regressive Foundation Model for Image Understanding and Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T04:34:11.061508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:34:11.061508Z digest=sha256:d84471d59955a41562bc4cfa4217e3f1e1ad8c7995927cebe6829435d96ce495

Observation 3e051434-6923-46ac-b071-f1409063ba81 · inbound

Transition Matching: Scalable and Flexible Generative Modeling cites this paper.

Transition Matching: Scalable and Flexible Generative Modeling Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T21:46:30.191271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:46:30.191271Z digest=sha256:ca35a39701b9daa4fd417af800368c5ec3252bcd8e1d7ab183b5604aab7f2cbe

Observation a3d45e71-80cd-47f3-bbb0-101445996b39 · inbound

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation cites this paper.

Beyond Patches: Global-aware Autoregressive Model for Multimodal Few-Shot Font Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-21T17:00:24.100992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T16:57:09.843691Z digest=sha256:0fb05a991a81bb9da9cb7447fd93555679fbf37180b525d96a67096be1b48cb8

Observation c1c71beb-6d0d-4eec-b642-46710fa28895 · inbound

Mirai: Autoregressive Visual Generation Needs Foresight cites this paper.

Mirai: Autoregressive Visual Generation Needs Foresight Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:17:52.072386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T12:16:16.461488Z digest=sha256:aa44e638f3f51c0d1039c53d9c438f0fa3d6cf9282095c8998084e3b2ee474d7

Observation a8c6eaa1-d52b-47f3-9230-edd866b78180 · inbound

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought cites this paper.

Shape of Thought: Progressive Object Assembly via Visual Chain-of-Thought Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T07:13:01.740855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:13:01.740855Z digest=sha256:bc9f664450308b121bf302265e507c8b1091b590929757c4737dcdb663875917

Observation 6c18d0c1-30cf-445e-9666-814e7aa69a2a · inbound

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models cites this paper.

Token by Token, Compromised: Backdoor Vulnerabilities in Unified Autoregressive Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:38:05.450573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T05:35:46.236860Z digest=sha256:ad085b3c8855150c09622f572f6d672a88d2fc6a84d8cacee3e5cc8f2cffb916

Observation eb19b34b-f87e-48c9-9918-9b17883db5ca · inbound

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation cites this paper.

FullFlow: Upgrading Text-to-Image Flow Matching Models for Bidirectional Vision--Language Generation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:44:03.082166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-21T07:40:36.206754Z digest=sha256:2110cb4b3c1c66634aee708830ff7db82d9fd9d550cb4e037862a4b7d70adffd

Observation 7533d770-7924-41f1-a666-d1af27ffe8ff · inbound

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning cites this paper.

InternVideo3: Agentify Foundation Models with Multimodal Contextual Reasoning Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 198

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:03.002018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T09:48:27.652901Z digest=sha256:4a353ef89635f456ae472962eb5bd0ecd00489ee4bb2a5a9f57e2dc0a206804a

Observation c5eac19c-a2da-47f4-9386-824529c178d0 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.661419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:123f64de8ab1dff42b485b9798370b284bafe8ecc9ef788f1bd9a9bf3890cefa

Observation b62c772c-45f8-4b51-b1fc-f3b5fa752002 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 70

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:19:02.540114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:0ffe8281d7f82e1f1be0ea952163d12c6d89cc73b7c902724193b986f36f4514

Observation 80d14c5e-5aca-41bd-bf26-dc472960ebc6 · inbound

Obliviate: Erasing Concepts from Autoregressive Image Generation Models cites this paper.

Obliviate: Erasing Concepts from Autoregressive Image Generation Models Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T16:35:50.633519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-30T00:28:37.837090Z digest=sha256:ce7a10489f2ce795612252f048756046e631ccfd85bee9fe54393b3cd5ce3b72

Observation a6b91e9c-04c0-486f-8eba-a1f81998cbda · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T15:55:25.318583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T15:55:25.318583Z digest=sha256:5bad3d3d3a405b621026bf89c566008dc3640568644440cdfe03337e8b75d387

Observation 6a7d5a16-0de4-403a-bb72-3732166dc2cb · inbound

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation cites this paper.

HELP: Human-Efficient Large-Scale Robot Post-Training with Rollout Segmentation Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T08:13:03.762013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:13:03.762013Z digest=sha256:138318e4a43647b28f66de20f6d4617376eb62c523f382b62cfc2dfde9605b8d

Observation 18ce830d-80e2-4c80-87b4-e5957605ca13 · inbound

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications cites this paper.

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T01:51:50.147250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:51:50.147250Z digest=sha256:41c9653d9fd245d0793e529cc89e7e66e980de6dbfb8ca686f46598b2fa0fbdf