Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:28:09.873139Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 34 inbound Pith citation observations for arXiv:2506.18095.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T23:28:09.873139Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T13:05:39.665891Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
69 of 69 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 69c5562d-0f13-4a78-8d26-0031208ab9d0 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b691a77-58c4-4ab2-90ab-953e6a7ba643 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Qwen2.5-VL Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c933785-3e08-43d6-b7d8-d0b435ead2e0 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Improving image generation with better captions
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70ff78fa-e686-41a1-8739-eadb3ae22e54 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unresolved cited work
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ef533ac3-2a70-43fc-b9e2-bd3923428b92 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbff37d4-5075-466a-a5d1-40fc08ab1094 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eac4d61d-d032-4427-a682-7161337aadda · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ALLaVA: Harnessing GPT4V-Synthesized Data for Lite Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c30384a4-4ee4-4ed4-a4e7-a5cea2b3ee23 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textdiffuser: Diffusion models as text painters
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a7146d27-7331-49b9-aca6-3cc16e2da07e · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8da046a-4135-48c8-be05-d26a0e3795d7 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85473dbf-336f-4058-a13a-706dc983d539 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2aaaca21-4a0c-4264-82cf-4fb231c583d6 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Towards injecting medical visual knowledge into multimodal llms at scale
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 957fd8ea-0870-43da-8c62-b3105c881098 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation An Empirical Study of GPT-4o Image Generation Capabilities
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d08889de-d609-4182-bf83-dc48d97d6027 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15442cd7-75d2-4569-9597-4fe2acd1e963 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 40d29074-f65e-410f-a2d0-d123bed6d857 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Imagenet: A large-scale hierarchical image database
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb9e5f82-cff9-40ff-96da-a3b628019a21 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Taming transformers for high-resolution image synthesis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4c5993d-a8bc-4bfc-8321-bf08e951258a · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Rectified Flow Transformers for High-Resolution Image Synthesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6433a298-c677-49e5-a1fd-37950c61de70 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Vitron: A unified pixel-level vision llm for understanding, generating, segmenting, editing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 87b0722d-66c5-4169-8e41-392d3cf70911 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f1804a4-f939-48ba-b28c-95fd060943be · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b44d41eb-f8f1-48f9-9326-50c82c9e473c · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6024e104-6b00-4937-8047-8d428796140e · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Laws for Autoregressive Generative Modeling
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f246f47c-5b94-48ad-9136-c22984658118 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Classifier-Free Diffusion Guidance
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38cfac70-e405-4248-a732-bbf13ee34dc9 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Denoising diffusion probabilistic models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4e2ec45-017a-4238-a3ab-4b06efe22c14 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06002829-b5c3-45a0-afbb-b23882b30f20 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation HQ-Edit: A High-Quality Dataset for Instruction-based Image Editing
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b25588c-f228-4a12-ba65-2cef87527286 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ba124ae-8510-4cb4-890d-62a2b53796e0 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e562b3fb-f63d-4cc7-b413-3d1c74b0aa45 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Microsoft coco: Common objects in context
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c300218b-9319-49f5-a65d-ccb32c38d7c8 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cff1c216-f4b1-4c03-b46f-4de3610ba0e9 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Visual instruction tuning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc9bb83e-0a06-4809-bc1f-7394afe78b75 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Step1X-Edit: A Practical Framework for General Image Editing
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c72b30d-b76f-46e3-904b-5b40ce4cc75e · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 274d2a4c-60dc-47e0-afae-701416afd088 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Introducing 4o image generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb7f369c-0901-4b24-9a8d-7f61448e8fda · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hello gpt-4o, 2024
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1f9c8120-20d2-4887-b0f5-f3a4ccebf895 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 833463f7-e50d-4d93-a8f3-1c9cb36c4c79 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Zero-shot text-to-image generation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72909ad8-c2ee-4c84-bf18-a35670526051 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c462b7d-8688-40c6-a78f-b0e484c67518 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation High-resolution image synthesis with latent diffusion models
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12a8bce6-abcc-4ab9-9525-50ad057c2fc2 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Laion-5b: An open large-scale dataset for training next generation image-text models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 222979f4-59dd-4405-ae40-8aaa9832f1ab · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textcaps: a dataset for image captioning with reading comprehension
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 51f3cb0c-7788-42fd-addc-18d22233d511 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Generative modeling by estimating gradients of the data distribution
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3edbe10a-0c9e-48bd-8f73-267640380da4 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c691b7-ca7e-43f3-9127-bcf2e679a695 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Emu: Generative Pretraining in Multimodality
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2dc2524e-c15d-4508-9692-c276ee0a6476 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Multimodal Latent Language Modeling with Next-Token Diffusion
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdd28907-eb26-4f3d-9ab0-5626dd9ad9f3 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c17f4fa-c2bd-4259-9522-c95de7b73cfc · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Gemini: A Family of Highly Capable Multimodal Models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c6b19f1-d144-4ac2-baae-41fc252e9ec4 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c624ce04-be73-430f-ac7c-a92fdeb31ee0 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyText: Multilingual Visual Text Generation And Editing
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50751418-9b2e-4298-80b4-0b39d30c5a8c · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Textatlas5m: A large-scale dataset for dense text image generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf8346c1-e658-4964-a6ab-e8e635871ba8 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Emu3: Next-Token Prediction is All You Need
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfdb897c-4045-4ef7-a3e9-6674083af233 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 495838a0-7edd-4c79-a3d1-63b7286c119d · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Next-gpt: Any-to-any multimodal llm
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ddc754d1-d7c9-41eb-957b-146175a044f5 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c02d195-e9af-4ad9-9409-7113542331ad · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation GPT-ImgEval: A Comprehensive Benchmark for Diagnosing GPT4o in Image Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0742d445-1c92-44cc-84ee-7aa31c60195b · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78361dec-7c30-48e3-bd4c-a52ff3363c04 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a485763-c51a-4191-a88b-d9b7a0856edf · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyEdit: Mastering Unified High-Quality Image Editing for Any Idea
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a224390-44a0-45a7-8793-be16ff421cfa · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation AnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d8274e0-8a91-441b-a62f-d36fd59cca34 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Nexus-Gen: Unified Image Understanding, Generation, and Editing via Prefilled Autoregression in Shared Embedding Space
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1116264d-5982-416e-b4cb-21b2edcf599a · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Magicbrush: A manually annotated dataset for instruction-guided image editing
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6a1d66f3-ab3f-4fee-89fb-2bce5ce3dbf8 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Ultraedit: Instruction-based fine-grained image editing at scale
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bfb47594-c57c-4467-b740-9912eabe5eb3 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8b3b05-e52d-4294-b1d1-4cad0a13b504 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd4dc842-09ed-4014-b8b8-caa0e77bb21a · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation write newline
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1018978-483a-44b2-8dad-4850576be72d · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation @esa (Ref
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f828fd-6031-4f3a-8220-d88fd0f0b425 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation Unresolved cited work
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da56fdde-9320-4ccc-9888-19350a643d65 · outbound
ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation v6kL8 | c6 +zP? z ۷ 7?Y| -w?iiiEޔCT t1gΜjsss + B M ]UVڡth<3 E NUݺu WM .\(P& @ס^^^= ;+[i4o ԗ/_FuPg- >ƹs ݴ_4
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 309f22f9-9241-4f0d-b7d2-f8eccc15f32e · inbound
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a4601e-976b-4aee-a89a-3b446283c17c · inbound
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5975655a-e5c2-4bd4-8859-fb51ab03caff · inbound
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9761af27-5252-44df-9dd5-1a00c12bafb7 · inbound
Reconstruction Alignment Improves Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc41ad36-ad34-4c40-83e6-3745a1c768ae · inbound
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5aa384f7-4c39-4008-bbee-35e3365c1c77 · inbound
Emu3.5: Native Multimodal Models are World Learners ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8a09de47-3500-4297-b64b-a42980b220fa · inbound
Distribution Matching Distillation Meets Reinforcement Learning ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 754a545d-74ad-4438-a302-90c6c1f46be8 · inbound
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cb519f03-a043-4030-8e65-b652ae1f883a · inbound
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa3362e7-21a0-402f-b312-f19d70ea8568 · inbound
PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e2391c0d-6086-4f42-b788-8402b6c26b5d · inbound
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f36e7ec0-3247-4d13-aab6-fb00b16848f4 · inbound
Automating Crash Diagram Generation Using Vision-Language Models: A Case Study on Multi-Lane Roundabouts ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4837438a-6459-479e-a0b3-ecbd2e941b85 · inbound
IncreFA: Breaking the Static Wall of Generative Model Attribution ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 09230663-92aa-4ccf-a58c-26893973f84b · inbound
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d9d9b23f-5dde-4f85-93ee-e9369d2e2083 · inbound
Guiding Distribution Matching Distillation with Gradient-Based Reinforcement Learning ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 036cb635-679c-45df-b969-4de48f914a80 · inbound
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation bb07873a-e4b5-4054-b07a-f76fb25daa70 · inbound
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 316fed8e-229c-45ae-93e4-189f4ff73509 · inbound
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0c7faca6-810c-47fd-a7bb-81d967daa41a · inbound
FlashAR: Efficient Post-Training Acceleration for Autoregressive Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6a7db49-d335-4b4a-810a-1172fdb05aef · inbound
Inline Critic Steers Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1766dc85-7e6e-424f-9f0d-7b44974e8ed6 · inbound
NanoCP: Request-Level Dynamic Context Parallelism for Data-Expert Parallel Decoding ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e6560117-548a-4d13-be28-0661929cf818 · inbound
Bernini: Latent Semantic Planning for Video Diffusion ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 85c82c7e-a6da-4d5c-8a6c-2801a4224f68 · inbound
Reinforcing Few-step Generators via Reward-Tilted Distribution Matching ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b5f23702-f0d4-4c15-b8c3-f3327b84f720 · inbound
Imagine Before You Draw: Visual Prompt Engineering for Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c6f74ddb-5aa5-4c90-aa36-b29908259095 · inbound
Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0b73f3c-addd-4bdd-8dfc-27ba463084f6 · inbound
ARM: An AutoRegressive Large Multimodal Model with Unified Discrete Representations ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ab01bb93-776e-4a24-beeb-28189b032601 · inbound
Unified Multimodal Autoregressive Modeling with Shared Context-Visual Tokenizer is Key to Unification ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe67454a-814a-44fd-ae3c-dd1fc17f989f · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1de9217b-9bef-4351-9d86-57918e098cee · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 641635b8-ad85-4a14-92e2-7750f4291c44 · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 103e9039-0169-4e0a-ab9d-31bb526dea8c · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ecef6b90-21a0-4183-9f66-858677c0af61 · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddcc576d-155f-4567-ac5a-5e08fe62b0e1 · inbound
Bridging Video Understanding and Generation in a Unified Framework ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7c2e7998-274a-41b1-a194-d37e832b4a2d · inbound
Amortized Moment Matching for Visual Generation ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation
Reference 150
Source-reported events for the cited work
Unavailable: canonical work link unavailable.