Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T07:47:34.464711Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 100 inbound Pith citation observations for arXiv:2506.18871.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T07:47:34.464711Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:44.468825Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-10T01:46:41.215139Z
93 of 93 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6a2ba722-9dcd-40fe-b431-f5efc73c9469 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ed1dec8e-fa1c-403f-b803-b5bd41cdf07c · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Sd3-medium
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 37fba2b1-a5d3-4d9b-a356-2f1065c21f63 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a316522-0901-470f-ab4c-bfcebb68372e · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Instructpix2pix: Learning to follow image editing instructions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a1e9f146-a17c-4674-acbe-9901027f0658 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Emerging properties in self-supervised vision transformers
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b9574b19-9b7d-4892-8a32-2e3c24f27396 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Allava: Harnessing gpt4v- synthesized data for a lite vision-language model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b5c5196d-3cf4-415e-9f1e-1c6775dacd86 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 363f5538-7668-43b2-a53a-bfa597dc22ba · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b7cb861-bde0-41f6-af84-380adb360b9f · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 459627b6-565f-4da6-87de-5c23393a4721 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Unireal: Universal image generation and editing via learning real-world dynamics
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9561aea9-5c12-4061-8594-2ce411eda1b5 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b6974b9-3ea9-4c9a-a6b3-2ecd153134b4 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5e74ec7-504f-4a6d-93b9-345b4de00768 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Emerging Properties in Unified Multimodal Pretraining
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fa2b8a6e-77da-4bcc-bbff-be3e0c57eb34 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Autoregressive Video Generation without Vector Quantization
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f17339e-d4e2-4030-8eb1-24b3a1c095a5 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 17db3860-bcc6-4b4e-92d0-b9e63ea08c72 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Doubao-1.5-pro
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 73068fbd-b414-4c55-9eed-2adb0db3fbe6 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Scaling rectified flow transform- ers for high-resolution image synthesis
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d162f4ea-428f-4dc8-bdf9-4b41940c6023 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation StyleShot: A Snapshot on Any Style
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d0872318-8561-4842-8a29-78fd6d1d6352 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9f47869-59a3-4b35-b475-99d7b8003557 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation be91d027-391b-400a-90ca-59b88100c947 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8c9a5eb4-2153-4a3d-a6eb-b5d005aade1f · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Gemini 2.0 flash
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 40cacb13-c58e-4743-8881-55e15c319a66 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b8fadb8b-3881-4257-9666-e71df3074dc1 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b6c0150-7cb3-40cd-bb14-3cf56b179ce2 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation MetaMorph: Learning Universal Controllers with Transformers
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 708fe622-ed8f-40d1-ac6c-1e5fb8996cb9 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7ae6e3b1-e8bd-410b-9bfb-824aa8e4a1c2 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Imagen 3
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4f53dc03-fe57-48f1-94b1-6ffcf0e6fefa · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation OpenAI o1 System Card
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8fdf8c54-6038-4419-bf5a-81cc3a9e0dd2 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 635951f3-409c-4028-b4db-369ceec4e30b · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c848a64a-5d7d-428c-8538-d435197390c0 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Auto-Encoding Variational Bayes
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e448e0a-1e42-49bf-9100-bb0aaddfd337 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Viescore: Towards explainable metrics for conditional image synthesis evaluation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c1dd45a5-f2ae-431e-9496-c7588126cbb0 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8ba481ab-2481-49ce-a696-8e57b8280e46 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Flux.1 kontext: Flow matching for in-context image generation and editing in latent space
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ff7cdb58-8803-421f-890f-df53b2a837c3 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2447e8fb-e23b-4880-adf1-7cbd280a3a10 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ea64e559-00bb-49fa-9315-8e5b1445afbb · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c9305dfa-6c32-4c91-ab1e-c02c83693a8d · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4e27def9-40d1-425c-9ce7-8b8a0ccec063 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 86a9cfa2-6016-473d-a430-11fa150207ec · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Let's Verify Step by Step
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 93ec51ed-148c-4c2e-aeaf-d570a1dcdf04 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 870aebc2-8f34-46cc-b66c-a74cf12c5759 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d29d5c62-c28d-4adf-86f0-cd93aa7ed8a1 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Visual instruction tuning.Advances in neural information processing systems, 36
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2662eff5-4533-4949-8b73-f916bcd7ddae · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 124f0a89-e720-4cb2-9bd2-75bd48985349 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Step1X-Edit: A Practical Framework for General Image Editing
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e121bcf6-50a6-4fa8-b5c6-1a481065528c · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4957671e-69c6-4a63-9541-755157b12c89 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 65d6bae9-8c2e-4b4f-91b2-dc18d2390135 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c076fba8-d99e-4b99-a298-f412ffb53052 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation DOCCI: Descriptions of Connected and Contrasting Images
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1c844a54-bd19-492e-91f6-0cc3cb9a68b5 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Dall·e 3
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ab9af409-c620-4a1d-8ca1-151328f48f6a · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bf2497c1-0f7a-43a7-b7ba-57d01c73db30 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 78d241c5-bccb-4740-99ee-1162d93747af · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation DINOv2: Learning Robust Visual Features without Supervision
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe261762-9aac-4b65-a482-88f113dac9da · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Transfer between Modalities with MetaQueries
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 346999b8-e40c-4e58-a75e-11e9d2767bf5 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61d4ba0d-15e2-42f5-8447-be27274c680b · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e14af303-202d-4ade-a62e-baaa88f27e73 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Tokenflow: Unified image tokenizer for multimodal understanding and generation
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eb303210-aa34-4c7e-9d37-c9c9a9640560 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Learning transferable visual models from natural language supervision
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d592384b-0c7b-4204-acb3-6fec65904c9e · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1d456f56-1889-441b-ba6b-3183b8b6d3bf · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation SAM 2: Segment Anything in Images and Videos
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 91190551-509c-439d-b59b-fd85a7951d14 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation High- resolution image synthesis with latent diffusion models
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8b33ffe8-3d3b-4a81-b14e-f01854e1f3e0 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5a859f6b-dc4e-4519-90e8-67c07c6d8e2e · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Laion- 5b: An open large-scale dataset for training next generation image-text models
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0cb7f001-34d7-467a-91e1-8031047eb52f · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Emu edit: Precise image editing via recognition and generation tasks
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 301bff5f-135a-4629-bd06-220f1bcd1fea · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 8d17d2e5-31b4-407b-b35b-1856094b0138 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Journeydb: A benchmark for generative image understanding
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 6d0613b1-39d0-4197-bcba-1f0903ca05ee · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Generative multimodal models are in-context learners
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 093b651b-a71c-4722-8a66-5191e621d724 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation OminiControl: Minimal and Universal Control for Diffusion Transformer
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f34466e1-4977-4e60-b676-d991eeaa3200 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 42930488-b5c8-4285-9f6d-7ae72bf657f4 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e5ba5c38-9473-4601-832d-f72cf168f5e3 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 77c7b205-9491-40d8-b778-e66da0da3f59 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c858461c-f18c-48d2-8d98-eae8801f644e · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Emu3: Next-Token Prediction is All You Need
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 140f44c9-653f-4573-926d-dcca532e1a19 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Omniedit: Building image editing generalist models through specialist supervision
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e755530c-fa21-4c72-aba5-6dfdb36e33d4 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ef199316-fc17-48f4-94e0-96ea07d976d4 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc61e83b-2629-4f0f-869e-418c538adfcd · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Omnigen: Unified image generation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 106f55b8-8ca1-4360-b981-428e9da0ba43 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b2798405-bc40-48f6-964c-ef41567dfdaf · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bbaac2e6-aa06-436b-81cb-af82f727f2dd · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c70239cd-dae7-452a-ae2f-665b7d5562eb · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Anyedit: Mastering unified high-quality image editing for any idea
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16e7f1ba-261a-4908-a691-8ec2a6e9b962 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5a54885-27f7-4021-b985-ff7c68482c12 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation PromptFix: You Prompt and We Fix the Photo
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 05ae6f34-1b01-423d-9f86-eb0ea44193e1 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36226127-c661-4c1c-9528-24bf2e23dab3 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Magicbrush: A manually annotated dataset for instruction-guided image editing
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 28d57a9f-f04c-4bf8-87c7-aa11c44bbb4b · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Adding conditional control to text-to-image diffusion models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16cd7688-c74e-478d-a161-de57d9098dc7 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ae88a07b-442c-4b3d-b86c-e65c9e7d154a · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Ultraedit: Instruction-based fine-grained image editing at scale
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation db41be30-8739-422d-9c6b-05007cc19616 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Uni-controlnet: All-in-one control to text-to-image diffusion models
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation dcfbb708-e06d-40af-b6a7-5c90929349a4 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71813034-7ca3-40fa-bc61-c7a1bcc69681 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation Lumina-next: Making lumina-t2x stronger and faster with next-dit
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eacc22c5-403a-4eda-8951-f1671aaa3d80 · outbound
OmniGen2: Towards Instruction-Aligned Multimodal Generation From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c67b9732-02e1-478b-8c54-be1a6568717f · inbound
OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57d336e4-0130-43a1-9f55-007a8c728c3d · inbound
Show-o2: Improved Native Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 117
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c6695ebe-3967-4611-900a-7da26f6548a4 · inbound
XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df0d8873-5fd2-4ce5-9671-a01b95f3a8e0 · inbound
Ovis-U1 Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74e7d532-c840-4003-b8b1-7aa5177ed517 · inbound
GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0317847f-23ca-44d5-ad3c-a9b2990ae033 · inbound
X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fcca5d7-f0e3-42cd-aa47-7cb2d58663e5 · inbound
Qwen-Image Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 276f6e35-9182-4dfc-a0df-9f322ea10bf2 · inbound
Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bac6495b-4c04-458e-b93f-75e5e925cdec · inbound
USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33e764d5-9f95-4abd-a8a4-42e0097ae95d · inbound
FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cd68c83-9b9a-4dc1-8953-e73c63c7c53e · inbound
MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5be15ec2-2c39-4bb6-ad8c-3a652241a015 · inbound
UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84579299-7ced-436f-8ae8-b816cbe7fedd · inbound
Interleaving Reasoning for Better Text-to-Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8339fa3-58e9-4268-a468-04193c791a09 · inbound
Reconstruction Alignment Improves Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41db8f31-f4b4-4a70-b193-2734420b0857 · inbound
FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4604da4f-afdc-492d-8ad8-bf46b607764f · inbound
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68362cee-3233-4960-8091-bf7eb83ddc59 · inbound
Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0d8e164-b510-4ecb-9384-2566a2feea1a · inbound
Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca619714-e36c-48d0-939b-d8e4c6ae6da9 · inbound
Adversarial Concept Distillation for One-Step Diffusion Personalization OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 32282b4f-8d65-4661-b3fc-7d3bb3cb4a24 · inbound
Emu3.5: Native Multimodal Models are World Learners OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 107
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9ce26491-9f90-47a9-bab3-09760e8f20b4 · inbound
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7491b3a8-d985-448c-a654-e3d6fca166e1 · inbound
DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fe4599e1-f29a-4df5-9c6c-ec9a0030f03f · inbound
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b91c64-a878-487a-bb71-a56fb007dc49 · inbound
AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bfe0fc3c-fa9b-4878-8e7e-c1fad62ebae5 · inbound
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 51eaf0c8-952c-4956-a7cf-6ffc95dbe6bf · inbound
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfdec00f-77ff-480a-b1b4-e4cb8b5927e5 · inbound
PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4132ca51-24ac-4cb3-a6ad-6cece81a4d18 · inbound
Reversible Inversion for Training-Free Exemplar-guided Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de38a0ad-99a0-4b45-b251-06a9f1714c97 · inbound
LongCat-Image Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9b43ddf8-2a52-4ed1-8c08-51624a0e48d7 · inbound
IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aeec8005-b4ab-453b-a695-d4ce9600be17 · inbound
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ec8b6838-caf7-4b18-a516-544778224030 · inbound
Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fa80018-870b-444a-9a83-36ae5772908a · inbound
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3077067-172e-481c-8614-c371964cb528 · inbound
Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d46e0128-8174-407f-a488-ffc1ec27817a · inbound
InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 892b9c55-ffb1-41ba-950c-15e9530cb652 · inbound
EmoCtrl: Controllable Emotional Image Content Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0fde9377-a7af-4fc8-886a-23b5b077b765 · inbound
A Unified and Controllable Framework for Layered Image Generation with Visual Effects OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 318707f5-890d-4c51-8954-70b29ae56e3f · inbound
Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f3f31d1-ceda-44ce-be33-d2661a195fb3 · inbound
OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cb75bae4-4c56-4e74-afd4-77cfa223ab65 · inbound
DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a3bdcf0c-3a3a-4055-8e2e-283613f01ae2 · inbound
WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a62bb644-a175-40c0-8495-aecb17fb2300 · inbound
VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d07ef0a-c423-43c0-bf56-8ca976ea3dbe · inbound
Demystifying Video Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9df42cb4-f0c1-4cdf-9ea2-3e0d77a4cf38 · inbound
Demystifying Video Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2debca2-c864-42a6-b137-d91ef1b07d80 · inbound
FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ba346ef-4557-4dbc-ac8e-74b445b4a9fd · inbound
TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d3ce28-6fcd-45cc-a024-537f7821a90e · inbound
HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 497b72ab-8d42-40fb-a74e-e3d27832cac5 · inbound
SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 120fb9cc-33fa-4170-98a9-9a6707b7fd14 · inbound
RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 97abfa03-6e42-4d1d-ae4c-c3913a7cd11f · inbound
InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e80d05a-68f4-4e28-8f5f-e15db5bd8618 · inbound
AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c25219df-fce1-47f4-a4f6-85975d722c22 · inbound
Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6778d0b-6094-4bd0-98ef-118084e33d7f · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cd658068-bc06-4582-8d91-32ef9fdc7a1b · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 882b8cdf-cdc2-4754-bab8-c7a7f7572e38 · inbound
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 396e7ff4-29e2-4f81-bbce-b73a27080530 · inbound
FineEdit: Fine-Grained Image Edit with Bounding Box Guidance OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3e29a58a-36e7-4657-ac6b-cb29f6e018fd · inbound
Nucleus-Image: Sparse MoE for Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29eab9d4-3139-40bb-90f7-97473d3f0d70 · inbound
ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 64c77953-466f-41b8-b759-dd3fed8388e5 · inbound
OneHOI: Unifying Human-Object Interaction Generation and Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 15f87b3d-ad47-4405-aa46-31afaf05db28 · inbound
From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 81307ca1-9661-4a25-b07e-3e8dae39c73f · inbound
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc121b08-a3b3-4fc8-add9-ff73d89f8085 · inbound
DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0987f470-dcea-4053-bc93-9ada02b2f8f1 · inbound
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6ea9339-fa88-4e79-b775-9971c7981965 · inbound
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25b8428d-e5c7-43d6-a6df-2f329507b3be · inbound
UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5b66350-6d1b-44ee-9cad-799389f02de2 · inbound
UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 29da7627-6712-49aa-9997-060008403a77 · inbound
HP-Edit: A Human-Preference Post-Training Framework for Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f412204a-5539-4ddf-a03f-eb4d327c7cea · inbound
SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 25659def-5c34-4524-a3e1-4dbaff10dd67 · inbound
MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0f6e9366-bb90-44f4-b33d-c01b29805084 · inbound
Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4bb6c475-8721-4cde-92b0-1ecc64bc29cf · inbound
Exploring Spatial Intelligence from a Generative Perspective OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b6c6055d-cb62-4252-b193-0b64a05811ef · inbound
Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 36dc96e8-3034-440d-9ee5-ea5454ba3ee7 · inbound
Meta-CoT: Enhancing Granularity and Generalization in Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71618ad8-1a68-4fc2-be81-02083c458482 · inbound
Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ee52f3b7-b9ad-4781-99ba-833876aff080 · inbound
DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d6b43f2f-f4f1-4bbf-8eb3-638ab79c0261 · inbound
SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b22ca79e-6e97-4824-9f5e-b28743a40476 · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ca22cb71-1ee9-484d-ae73-97c5cd52e08f · inbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7d376e86-6a5e-4d49-9e24-016473c1fe71 · inbound
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 145
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation eff83de5-2056-4c0d-ad4f-227e41458c57 · inbound
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 71411b96-9953-4996-98b8-e2b3ee92420b · inbound
InsHuman: Towards Natural and Identity-Preserving Human Insertion OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 192fc214-3048-4529-bc50-b3b36f17304a · inbound
EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c532411a-81bf-42ba-93d7-3a80dde2550b · inbound
ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cc490d0a-6be6-42ef-84aa-6f2a880b22d6 · inbound
STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7234789c-dc72-40f9-ad46-ef7ee8317a99 · inbound
MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 74ef5eb8-a994-4b05-96d2-987a0ca67840 · inbound
Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 59881774-674f-4618-88c4-c226a47c88a1 · inbound
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c729739c-b642-4803-9bbc-15de28edd2e2 · inbound
Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b79e1881-614d-4aac-8dda-75c2ab16b2c5 · inbound
Masked Generative Transformer Is What You Need for Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cdfc8470-3fbd-4dda-8c09-5149de8c4930 · inbound
HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f4a40299-e364-4a17-8978-75fd63230e32 · inbound
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 39636689-405e-4099-bc0f-b71c36f99636 · inbound
RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 16d0b599-25fd-4eda-a5df-3676c5a9686c · inbound
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f193d64a-74da-4eac-ae58-191a2e9da181 · inbound
UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0b1bfc16-6493-483a-810d-131246f4bee9 · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e9c6ffa5-2c1e-4d62-8035-7e5b067cd766 · inbound
Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4f78e979-c64b-43cb-990a-aa9b56877b56 · inbound
Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 80bb53d2-99c2-447a-a49d-f0a9bd9b621e · inbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e01e0ef2-ea63-403a-82d0-c19828723b73 · inbound
Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4b2c1e66-1d52-489a-9c4b-96e5a9e491a3 · inbound
Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.