Pith. sign in

Paper Citation Record · LEDGER

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

As of 2 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 1 inbound Pith citation observation for arXiv:2605.12305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12305 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:48:04.997796Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-02T06:30:47.504484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:27:25.307245Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact32
  • verified fuzzy17
  • unresolved0
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 184fb529-11ff-4f91-bc77-d47362a88a87 · outbound

This paper cites Qwen2.5-VL Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.763962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:26a8613225a272dae9ae920625f5a1e6cf23232a54ca73e76edef33e43fb5cfc

Observation d02cc122-6a9e-45fd-b276-cb6cb638b2e9 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emerging properties in self-supervised vision transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.861704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:280f5cbd23622e4a6047ffcf915dc21c4a67b425e761ae13054d38587868168f

Observation 34142fe9-7a68-445d-94dd-dc4d9f9e8886 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.690835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:19c858a3a43a125df4c9e6c9d2b99959ddb78c2d29b99333404ba9656898cff4

Observation ec5f7b26-8613-497c-9838-d9e2a579dbb4 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.743962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:ef52290c2664b978bf47f32e53295cb6da5eb93ce16a38b4740b5ce8265a15e6

Observation a77391f2-5a52-42be-a84f-9723756559f7 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.720119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8883105aad4085d0b091e189319c37390c2610b1ac6b7540bf56d992531bfa97

Observation 6859090a-7d47-493b-8db7-75ed1d5ad84a · outbound

This paper cites Gemini 2.5 flash image.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gemini 2.5 flash image

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.807997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8f98ca6dd44db9878ab342008199f398d5925a168a746933e40c544467e3337b

Observation bde4c4cb-f641-4bc1-8f6e-8da44201c409 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.754291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:4583c76045ec54c80742ed9a9dafa3d366c6f06d681b1b23dd409d25e1ebaf51

Observation c15bc8fe-bdfd-4bf4-8cdd-171527205dce · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.814309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:0e7f06c0a29e1a84a14cf147571bb963326c620dced9d09693634a76d5d31594

Observation a915e8f3-18a6-447e-9509-ea2dccc0f5aa · outbound

This paper cites Seed1.5-VL Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seed1.5-VL Technical Report

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.697607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:5277b792bbc073fcddd79fcf7599635edf01fbd53dab64d720747ef78fa98ff8

Observation 95701d1e-0e24-4826-bd44-f3772ff1fe0c · outbound

This paper cites Chameleon: Hierarchical clustering using dynamic modeling.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Chameleon: Hierarchical clustering using dynamic modeling

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.842703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:5bbf2b19e5f42028359fab5d84d681e98a2831b4f64dade666320673f31e030f

Observation f8dd4f0a-f458-4477-b5b8-369d13a354ed · outbound

This paper cites Segment anything.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Segment anything

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.803268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8cd7a22b438a55c95688d84c7e06dd960e77c12d3d142db92583f0196bb5f9fc

Observation a41534cd-f50b-4f6e-9449-24649a562a6c · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Flux.https://github.com/black-forest-labs/flux

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.794238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b3a626d0e06ec7aa05b5828395272eb985397c37ece1e04d0f00dcfed9514202

Observation 691e6ee3-268f-4a8a-85c2-3db6d39c2708 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.757588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e4e7846b218bdd41d16668ae31d8581619e23a30db7b8787a7a70aaf6727079f

Observation ed4986ff-549e-4aee-8c4a-8a21cccce593 · outbound

This paper cites OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniCorpus: A Unified Multimodal Corpus of 10 Billion-Level Images Interleaved with Text

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.782843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8e7e2ec8dd96e88faee6b067b1dae0daf81490bb904344a6e0daf0ea31552904

Observation 87b58bcf-d74a-4665-b981-7d4557b10129 · outbound

This paper cites Describe Anything: Detailed Localized Image and Video Captioning.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Describe Anything: Detailed Localized Image and Video Captioning

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.777196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:5e22c7b5009135727072ec6c6626c76a3ef9bcac10264cb4fe0f5cc814d5ad5e

Observation 40d95ae1-9ffc-4cf6-8974-1b2b6cbd2541 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:9a7d0cb20728155740e5b744899374cca3cff3b5ef7cfe2830bc19c74a60fd0a

Observation d995ee84-a320-4df2-89ed-2763e9389dc6 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.785528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7cb598cf7c6dc20b57e95fd4bf99c784eedc98346c6bb8c5d3fcf18669ea06b9

Observation 6e416529-8161-44fd-8d2f-6e6778ffb2ef · outbound

This paper cites Visual Instruction Tuning.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Visual Instruction Tuning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.790888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7723025f960b11d1af67488740421c8d76d4e8783e7e41cb8b7568b9afb63547

Observation c6a591f9-b05a-4d3d-8321-b25e1edf1490 · outbound

This paper cites Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Lost in the middle: How language models use long contexts.Transactions of the Association for Computational Linguistics, 12:157–173

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.781426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:66b4d06bd6397f611a22df05b23065d7377b79eecae4b2f9ac8a39f4df1ff1f2

Observation cbf2701d-f968-4984-b387-5143eccc9a18 · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.788227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e6aa56535ad647f676ecd83c107ff96845bb8a6ab94b7cc783303d96ad9db123

Observation 79b11037-e058-406e-a1c6-9b940fedf1ec · outbound

This paper cites Dreamo: A unified framework for image customization.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Dreamo: A unified framework for image customization

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.773926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:710a3ee9db390e58133404f96ff5105892a27c6e98c435521ac90b2b2b7f5a46

Observation 0da8606e-f67f-4aae-85d1-6590d4590e00 · outbound

This paper cites GPT-4 Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation GPT-4 Technical Report

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.737992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:051f452a044a2ba050f67185a490e5e0520cebdb76f3afba9c43852c5e6ff347

Observation a830e551-45e3-4a76-807b-4c882d120aa2 · outbound

This paper cites Gpt-4v(ision) system card.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Gpt-4v(ision) system card

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.798877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6d1110aa9eb96deb5a60a8d515e766dd895deff181e8270ff67378efaab8f2ff

Observation 8df37de9-e8ca-4b22-ab98-0957be847c64 · outbound

This paper cites Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Introducing 4o image generation.https://openai.com/index/introducing-4o-image-generation/

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.847133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8a7399cbfabdeb6e2650698a7dfe8edb10e09f0cccbf2ccbc31c27bf0da556f0

Observation ca46de99-96da-4d03-9951-b21a76c762d2 · outbound

This paper cites an unresolved cited work.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Unresolved cited work

Reference 25

Resolution
parse uncertain
raw_fallback, observed 2026-05-13T10:12:37.851695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:1c448e39e14b1ef83f4f5fa531b562d6b60b67c036b58e0f5ab6a04a3ac61f0e

Observation 74bb8080-a6af-4291-9980-545ac4c94525 · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.770498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:ec004b89a2a81f4c0b5bf72c727aec0fa9ad2e03f8b588a23b36f14a52f78d22

Observation e6de04df-cdf7-4de2-bf8a-5f4e2345d248 · outbound

This paper cites DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation DreamBench++: A Human-Aligned Benchmark for Personalized Image Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.767075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b11b880b92c0dd301f9e17ec35f20bb7e7dca04293d3ca5354b0c137b6e6f2ce

Observation b9cc249d-4f56-47f3-9fad-7f0522c1758b · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.760735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:26aee5335a849e4946322ac9f222bc30e0512b24a76ebb01416635e08aaade2c

Observation 35079abb-46b9-499f-8a98-7330939ce200 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.710492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e87ce39d770890ffaeb1a6dcf9d9b9d5e2506204d367abd25d8df56ee7b2677e

Observation 29f1fd73-3ec2-40f1-831a-49f943a99bd9 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Learning transferable visual models from natural language supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.856835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:857f6ab2aedb8cdd72ccf96af46ac9fbda90489ed602694618de49e7ec5618ac

Observation c8fe1854-57a6-4bd3-be70-5172928e8262 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation High-resolution image synthesis with latent diffusion models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.869908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a3574f5f92981b53f503aee47c590ab1c7b6eaa15ee32f3812afab571b914f07

Observation 05a1c1d1-8de2-4676-b1fe-d306af5cb477 · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.701026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6881957cadfaa6c67ba892c1edbc99199ae8d52aa1551575f0bf3295564bfdf1

Observation d4776120-3d4a-4c3d-8792-798b1b8513d2 · outbound

This paper cites Generative multimodal models are in-context learners.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Generative multimodal models are in-context learners

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.865865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:9d36face092b672aa063955b8b11c0e8832fb739eb2f71206caaed94511ce8d4

Observation 815dab61-6538-43db-838b-ac5235468c31 · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.726708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a248504d510fd1904a61a16d95b30d05178f3ceeffcc1e6d5159e4a009e679d4

Observation 1a4ee54e-4b6d-44b7-b286-383a13866abd · outbound

This paper cites Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Reference 35

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:52:22.732596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a522b15e9c6c5681e2add5434357dac99094b6f92e885bb969444bfc08bdaede

Observation c181c187-ec0c-430d-af82-7e915d741b23 · outbound

This paper cites InstantID: Zero-shot Identity-Preserving Generation in Seconds.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation InstantID: Zero-shot Identity-Preserving Generation in Seconds

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:02:41.552159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:7b0476ae00647e82a610ae798866f3dedbcbe10254ce9d44ce1b809f958525d8

Observation 34e5f6b8-aa7f-488f-a273-e231abb719ad · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Emu3: Next-Token Prediction is All You Need

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.729608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:405b397b536f62f80473224c23e24a1e924092ee6be42496aacd01324d1254a0

Observation d30ce7dc-45e1-4262-abef-f8d354840325 · outbound

This paper cites Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.779971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:8fe7dc9adb66cb0afea6e9489cd4a04eaa1a4aa200b5ac52e0aff8cac13e3d75

Observation e44053e8-a5c4-4de8-b981-70520fe6bcf9 · outbound

This paper cites Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Elite: Encoding visual concepts into textual embeddings for customized text-to-image generation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.838638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:a61c40278f526df14764c68480a59ae712b59d3aaf7d153b8a6c911b0c70478a

Observation 460aecfc-05ad-4a34-a506-df0bed331fab · outbound

This paper cites Qwen-Image Technical Report.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Qwen-Image Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.722761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:811f081837bfd4ba8c1e701a42881636861dff43207d494afe9fdc193dbc9bfb

Observation 9ce679b5-b696-4cd5-a973-71cae6ab9f6d · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.833070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6c623d3888b81da613d983453c3ebb68efb9593240c383c6a9b8caf60158c800

Observation 4f78e979-c64b-43cb-990a-aa9b56877b56 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.735182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:e573bab116a8245e8f70a2750b81e2973273e6b096b56db39d7e7468930531fb

Observation b6af3b54-1388-4e76-b3c0-4b9ffa81aff5 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.714030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:51603f80501262815d917f1820ad670299d8a1ad6e4e9444fa6fbf7403e91c1b

Observation ed38d77d-eb02-489f-859a-713aecd61eca · outbound

This paper cites Dreamomni2: Multimodal instruction-based editing and generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Dreamomni2: Multimodal instruction-based editing and generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.750928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:3140a4dc98b3e0dfd27071487b5f04a49f56f3e26ab0e20e12f7ea7fd92964a2

Observation 8aa35615-b44f-4bbe-8c39-2e55c8afceb0 · outbound

This paper cites Omnigen: Unified image generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Omnigen: Unified image generation

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.825560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:b158dd7b8e9b2043bba1f246484a52e51a696bd6b70c40e27bf253276cd9dba5

Observation eb73f5d1-3058-4b12-84d0-33a396d6d6f0 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.707578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:d01a5db0f5e032b3cddd1539d4949a817cadde04820b5344f74e8234ef4a4dec

Observation 3c2fbb64-7882-4437-97cb-ce9e9ec09fd4 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Show-o2: Improved Native Unified Multimodal Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.747295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:1c846ac36b485f235aae16e93ca5e2049b8a9b462f1170c263b71beaa0ad484e

Observation ccbd9b75-0f48-4282-9465-0574330b9c6d · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.741046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:08fb408636e49868718797eec86304e2dad61c2ab9c42fb44284ca3d5522c01f

Observation 9f65ce3f-0c10-407a-ab3a-167a1e28387d · outbound

This paper cites Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.717435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6d4aa871449a675786ad6aaa66361f447d1975b6daec65d335446bc25a9e9afc

Observation facabc92-ffa3-4be5-9c3b-ec6978ae69b0 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:27.005672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:94293d919be1b7a674639e803838dd43cb3280159a6f143ea52e6cd7374eb061

Observation 81863fae-9bfc-4f09-891e-b9bcbbe42cf2 · outbound

This paper cites Image First.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation Image First

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:12:37.820894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-02T06:30:47.504484+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:6ffbc13851449b3900ce74775e4d0d68e0efa97a9b0b99282a252042f00de1b1

Pith citing papers

Observation 46ed4f89-c7f1-4cdd-abe3-920c94449164 · inbound

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships cites this paper.

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-02T01:27:25.307245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:27:25.307245Z digest=sha256:4736da48fa91b1f4658cb6b586c45efcb903c0db958fd23b0028b0e411682a52