Pith. sign in

Paper Citation Record · LEDGER

OmniGen2: Towards Instruction-Aligned Multimodal Generation

As of 11 August 2026, this Paper Citation Record lists 93 of 93 outbound references and 100 inbound Pith citation observations for arXiv:2506.18871.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18871 v4

Coverage vector

measured 93 of 93 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T07:47:34.464711Z

measured 193 of 193 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 100 of 168 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:44.468825Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T01:46:41.215139Z

Reference resolution

93 of 93 outbound references displayed

  • verified exact53
  • verified fuzzy35
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6a2ba722-9dcd-40fe-b431-f5efc73c9469 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.741429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:eb334e4d7049fde79dba03341768362b2eeb12219f65ab299ef1f11d4f6628f7

Observation ed1dec8e-fa1c-403f-b803-b5bd41cdf07c · outbound

This paper cites Sd3-medium.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Sd3-medium

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.270847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5e1266f1c8d0ad114c26ced851a8916610d011ad7984d77575b3d4b2373c20a8

Observation 37fba2b1-a5d3-4d9b-a356-2f1065c21f63 · outbound

This paper cites Qwen2.5-VL Technical Report.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.733093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:1d96ea4767069080b5ab3e8becd76c012aed0ccd799484649539829593e25ee1

Observation 5a316522-0901-470f-ab4c-bfcebb68372e · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Instructpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.274955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:adfc839d13559d9dd91d9e7faa9fe24a9a810f1c0cb958504a5caf28116f0e96

Observation a1e9f146-a17c-4674-acbe-9901027f0658 · outbound

This paper cites Emerging properties in self-supervised vision transformers.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Emerging properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.278398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b62f80f75332ee4ac7f4e12c8ab6e999f7573586efd4c1af335259cf33cfbb05

Observation b9574b19-9b7d-4892-8a32-2e3c24f27396 · outbound

This paper cites Allava: Harnessing gpt4v- synthesized data for a lite vision-language model.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Allava: Harnessing gpt4v- synthesized data for a lite vision-language model

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.267330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c95964e837848360746ee2fd4e843e354f7aaf4db9d334c04a702f2068368e9d

Observation b5c5196d-3cf4-415e-9f1e-1c6775dacd86 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

OmniGen2: Towards Instruction-Aligned Multimodal Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.737237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:301d7c2b3afa60aa940dfbd70370848ca4a6b91099e6a896d0b23052cf5c9a3e

Observation 363f5538-7668-43b2-a53a-bfa597dc22ba · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

OmniGen2: Towards Instruction-Aligned Multimodal Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.913871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a9b9636b37bbcf3d7378fa12758b51562ded00790f6189f0b2f5cf3b8897b185

Observation 8b7cb861-bde0-41f6-af84-380adb360b9f · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

OmniGen2: Towards Instruction-Aligned Multimodal Generation ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.855213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:745e2f302942fe83fe61e035f4920d3cd73778bda8c05622dfea28b99c5b7e6a

Observation 459627b6-565f-4da6-87de-5c23393a4721 · outbound

This paper cites Unireal: Universal image generation and editing via learning real-world dynamics.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Unireal: Universal image generation and editing via learning real-world dynamics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.400653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:22482952cd6dd574802caba585cd7ebb75516e11fb074be986c6af8fa249adc1

Observation 9561aea9-5c12-4061-8594-2ce411eda1b5 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.813557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b2d6c4eb8eff614e4f1dadc5cb00ebd7e063b0fc1ffda35b2bc2dc7434afea7e

Observation 9b6974b9-3ea9-4c9a-a6b3-2ecd153134b4 · outbound

This paper cites Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Davidsonian scene graph: Improving reliability in fine-grained evaluation for text-to-image generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.315466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:3e6914a69dc4d3d8af599e4353b5952405e7001481ea92de5451d6d665edb34f

Observation d5e74ec7-504f-4a6d-93b9-345b4de00768 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Emerging Properties in Unified Multimodal Pretraining

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.862080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:8b9e5671af85512bbcbb3c87d35730a28e51741596b7ed4dd3ee5a3a3feab4d8

Observation fa2b8a6e-77da-4bcc-bbff-be3e0c57eb34 · outbound

This paper cites Autoregressive Video Generation without Vector Quantization.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Autoregressive Video Generation without Vector Quantization

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.797784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:ede823b26b445d207de16be0acbce912f5571de622eec843f8a16fedc8018453

Observation 5f17339e-d4e2-4030-8eb1-24b3a1c095a5 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

OmniGen2: Towards Instruction-Aligned Multimodal Generation An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.784251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:8341e09389db6a28c93f9c8096d241030b1926c098368728662a89887ad69602

Observation 17db3860-bcc6-4b4e-92d0-b9e63ea08c72 · outbound

This paper cites Doubao-1.5-pro.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Doubao-1.5-pro

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.424980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:647d3933862c203d32e2a429248727567fad467770ca47cd21ee457044ca70e7

Observation 73068fbd-b414-4c55-9eed-2adb0db3fbe6 · outbound

This paper cites Scaling rectified flow transform- ers for high-resolution image synthesis.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Scaling rectified flow transform- ers for high-resolution image synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.395847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5ee48c66452caacfcb511a8966477c5b3d3c62dfb690f557973390bf9b721dfe

Observation d162f4ea-428f-4dc8-bdf9-4b41940c6023 · outbound

This paper cites StyleShot: A Snapshot on Any Style.

OmniGen2: Towards Instruction-Aligned Multimodal Generation StyleShot: A Snapshot on Any Style

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.849087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:e06c9b9997b12338e26b8575f9ea82bc22ef4aed8b7d4b901583da1e0654feb8

Observation d0872318-8561-4842-8a29-78fd6d1d6352 · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

OmniGen2: Towards Instruction-Aligned Multimodal Generation SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.777831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a399cce118a0d7e36292e0726a737625c5a582f1d8fd9fc85d858e4f469b5abb

Observation e9f47869-59a3-4b35-b475-99d7b8003557 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.394631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:bc9f833c87789958a300c09b7351b9967bec7a7758956de443bbd1f1bb97693d

Observation be91d027-391b-400a-90ca-59b88100c947 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.373080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0c49cac0fa505951c1a9625619a0d6606a4b271f7aa4f0e0ec2828e7a05cc2e8

Observation 8c9a5eb4-2153-4a3d-a6eb-b5d005aade1f · outbound

This paper cites Gemini 2.0 flash.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Gemini 2.0 flash

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.378218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0a5d336d003348c413135a6a2cc67a273d42eb964e57f14d095197cbcade58d8

Observation 40cacb13-c58e-4743-8881-55e15c319a66 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.752625Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b14cb0940ef145809834c8ebfb2707f3ecafe2cd15312ee0a5fa04964b639441

Observation b8fadb8b-3881-4257-9666-e71df3074dc1 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.744905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:248cfe84c85c15c335542c102955b20c8ef7927620b71a4d31f3e97c2dfd3802

Observation 9b6c0150-7cb3-40cd-bb14-3cf56b179ce2 · outbound

This paper cites MetaMorph: Learning Universal Controllers with Transformers.

OmniGen2: Towards Instruction-Aligned Multimodal Generation MetaMorph: Learning Universal Controllers with Transformers

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:52:10.748377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:aaa02cf175ef26d16323862661c079748db3770cbe8f5a9a50a4d64ea87ff05d

Observation 708fe622-ed8f-40d1-ac6c-1e5fb8996cb9 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

OmniGen2: Towards Instruction-Aligned Multimodal Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.756178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0f37a85b437a8165c7fe03554dc8bc8c77c99be771c205cdecdc1f7cbc88e591

Observation 7ae6e3b1-e8bd-410b-9bfb-824aa8e4a1c2 · outbound

This paper cites Imagen 3.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Imagen 3

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.767082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:2131f1a09ab650be3ac7f8c9d31fe53563286783884f39c321e55edf8c5ec9ca

Observation 4f53dc03-fe57-48f1-94b1-6ffcf0e6fefa · outbound

This paper cites OpenAI o1 System Card.

OmniGen2: Towards Instruction-Aligned Multimodal Generation OpenAI o1 System Card

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.852094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a447cb585629e84c3c70c7fd41912472521f4179ebfed8724fc7df09e7cc1bb9

Observation 8fdf8c54-6038-4419-bf5a-81cc3a9e0dd2 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

OmniGen2: Towards Instruction-Aligned Multimodal Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.955016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:efb24e45712b1ea399f3e1367d3e8df61ab019eeb065b6b8e368a426ed02970c

Observation 635951f3-409c-4028-b4db-369ceec4e30b · outbound

This paper cites InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity.

OmniGen2: Towards Instruction-Aligned Multimodal Generation InfiniteYou: Flexible Photo Recrafting While Preserving Your Identity

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.885914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:609069eba7c6d58f4e9cf1ab3fb6a3970e7d96249c5b0759859f3b81740caf3b

Observation c848a64a-5d7d-428c-8538-d435197390c0 · outbound

This paper cites Auto-Encoding Variational Bayes.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Auto-Encoding Variational Bayes

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.889589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:964d36587a4dc57516dd5a3f77d1312a6f58fdbc4880b82c9d60c6aa3e819e52

Observation 2e448e0a-1e42-49bf-9100-bb0aaddfd337 · outbound

This paper cites Viescore: Towards explainable metrics for conditional image synthesis evaluation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Viescore: Towards explainable metrics for conditional image synthesis evaluation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.355036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:26083b0f12bf3c55e6acebb1c6e5396c4d93f36de01b20bed161268a7dfd2c36

Observation c1dd45a5-f2ae-431e-9496-c7588126cbb0 · outbound

This paper cites an unresolved cited work.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:23:03.375056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:fcccc3058b03d778e04d1d3766d35a305d6cf8d51dc9a467360be4f403d8430e

Observation 8ba481ab-2481-49ce-a696-8e57b8280e46 · outbound

This paper cites Flux.1 kontext: Flow matching for in-context image generation and editing in latent space.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Flux.1 kontext: Flow matching for in-context image generation and editing in latent space

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.371597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c453433811e276da87002194659fef239d57fe7a10b016aa1b092fbd4fe635b3

Observation ff7cdb58-8803-421f-890f-df53b2a837c3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OmniGen2: Towards Instruction-Aligned Multimodal Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.882399Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:f2978a1230d7a0937fa01926795b01c7eb62461ede3c2ca638628cd5e65d94be

Observation 2447e8fb-e23b-4880-adf1-7cbd280a3a10 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.875408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:003dee7039c660a962ed03742de5cbc10425195e785b7ffc82bf75cf94e6f8e2

Observation 8f9c9d56-a50f-40e0-886d-c723d7ea8aed · outbound

This paper cites What If We Recaption Billions of Web Images with LLaMA-3?.

OmniGen2: Towards Instruction-Aligned Multimodal Generation What If We Recaption Billions of Web Images with LLaMA-3?

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.869347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:6e694159dc0515c5f12bf5ab743c2ee54a43ba8c0c4b56566875ea1493ca8628

Observation ea64e559-00bb-49fa-9315-8e5b1445afbb · outbound

This paper cites DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DenseFusion-1M: Merging Vision Experts for Comprehensive Multimodal Perception

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.872328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:39381454318dcd525bb2cc15f2bdff3122ec420e7166cc042c259300af593160

Observation c9305dfa-6c32-4c91-ab1e-c02c83693a8d · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.879158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:45c24f4018aa0fdc74a67959c0485177f2aa896e545037e6dd4771959415f555

Observation 4e27def9-40d1-425c-9ce7-8b8a0ccec063 · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.893022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0ede26670d7c1a90eb1be9c983192323d594a1250f7b8b7e77f32575ca8dbd4c

Observation 86a9cfa2-6016-473d-a430-11fa150207ec · outbound

This paper cites Let's Verify Step by Step.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Let's Verify Step by Step

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.961184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:34e727ec3971b298f18200caaa44f262300a11b19a19ecc8bf549078ba6b3b62

Observation 93ec51ed-148c-4c2e-aeaf-d570a1dcdf04 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.865617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:e8428d1e0529d958f4c9bace1e981c4d0918b719f52751277d14c6b765e1699c

Observation 870aebc2-8f34-46cc-b66c-a74cf12c5759 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.426447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c419f88189501e26b443777205cd31f81a833a4f7a3ced574ba86c714bb97950

Observation d29d5c62-c28d-4adf-86f0-cd93aa7ed8a1 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Visual instruction tuning.Advances in neural information processing systems, 36

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.346569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:761eb2dafe905164a6f022f9f9229fa233d1c0bdffeddad2cfd8f44b516331c0

Observation 2662eff5-4533-4949-8b73-f916bcd7ddae · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.804042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:f5c8ed1ed0fa0d899e8d821aee1fc92e3b1251fd3267ebebc987165d37918378

Observation 124f0a89-e720-4cb2-9bd2-75bd48985349 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Step1X-Edit: A Practical Framework for General Image Editing

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.845232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:4ab0e7ec1ae31ae8a191352fa22d4a1d260161761e8ce36feb6295bdba1bc79b

Observation e121bcf6-50a6-4fa8-b5c6-1a481065528c · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Mmbench: Is your multi-modal model an all-around player? In European conference on computer vision, pages 216–233

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.404652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c647fb9f8f5ad5da8282a2f37388f2a2bcee413f000bbbed20095dfdb23732dc

Observation 4957671e-69c6-4a63-9541-755157b12c89 · outbound

This paper cites ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling.

OmniGen2: Towards Instruction-Aligned Multimodal Generation ACE++: Instruction-Based Image Creation and Editing via Context-Aware Content Filling

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.791115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5fffc8ae0517b3e72557d9d6c3312c2d7a79efa738cc472ededb8223269ce70a

Observation 65d6bae9-8c2e-4b4f-91b2-dc18d2390135 · outbound

This paper cites T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation T2I-Adapter: Learning Adapters to Dig out More Controllable Ability for Text-to-Image Diffusion Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.800768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:32ab5b60687034d32a8d86a899096d3ce21852d5eb4ed1547f5f86e135fe4561

Observation c076fba8-d99e-4b99-a298-f412ffb53052 · outbound

This paper cites DOCCI: Descriptions of Connected and Contrasting Images.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DOCCI: Descriptions of Connected and Contrasting Images

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.787801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:91376a1f58ccc9c2b2f3b20a494d3cace74e74d1b535cc41a0470edd4a2cba56

Observation 1c844a54-bd19-492e-91f6-0cc3cb9a68b5 · outbound

This paper cites Dall·e 3.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Dall·e 3

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.334121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:85d0d9c17f627786b198dfe03fddee42e5e1733451d4e0e84e0ae055ccf02362

Observation ab9af409-c620-4a1d-8ca1-151328f48f6a · outbound

This paper cites an unresolved cited work.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:23:03.367588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:beb6163ff567cad85c70c0b8164a22bd43872f3aa86f8fbae1444c686a8fdc0a

Observation bf2497c1-0f7a-43a7-b7ba-57d01c73db30 · outbound

This paper cites an unresolved cited work.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-05-19T08:23:03.359229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:db24189ecb022becbb78bd1a93b8cdbda29f59448e3d90f49b76a6270e71b6d9

Observation 78d241c5-bccb-4740-99ee-1162d93747af · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

OmniGen2: Towards Instruction-Aligned Multimodal Generation DINOv2: Learning Robust Visual Features without Supervision

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.794403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:265f677e3182c0436778e18d731c86916d303291b92fc80345a1b5d341bd72c2

Observation fe261762-9aac-4b65-a482-88f113dac9da · outbound

This paper cites Transfer between Modalities with MetaQueries.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Transfer between Modalities with MetaQueries

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.770923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:9c911f30a0b322c7a9451713fca3a6399a9db8d496636c69dcfac47ffad8d5f4

Observation 346999b8-e40c-4e58-a75e-11e9d2767bf5 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

OmniGen2: Towards Instruction-Aligned Multimodal Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.780905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:ef866948f9d986bedd10f273655a9b23fbf1222c64e39ef57c32b74854e2be5c

Observation 61d4ba0d-15e2-42f5-8447-be27274c680b · outbound

This paper cites Lumina-Image 2.0: A Unified and Efficient Image Generative Framework.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Lumina-Image 2.0: A Unified and Efficient Image Generative Framework

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.760173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:08624bfb6889c89f57e2e8db4dfb910c66fba6b1fb5d396f3bc86933f639c266

Observation e14af303-202d-4ade-a62e-baaa88f27e73 · outbound

This paper cites Tokenflow: Unified image tokenizer for multimodal understanding and generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Tokenflow: Unified image tokenizer for multimodal understanding and generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.415719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:67e2dcf659c02e23d8b632641ef364071a6d6c6c5a284a286e71ad64d53f02ec

Observation eb303210-aa34-4c7e-9d37-c9c9a9640560 · outbound

This paper cites Learning transferable visual models from natural language supervision.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Learning transferable visual models from natural language supervision

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.369061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:fe7cf2a18933a78019d52c6f93d7b1312b78d0a385be60948db420639acf4ad3

Observation d592384b-0c7b-4204-acb3-6fec65904c9e · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.763365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5c2a29ffbab225432cb4fd2d6496bb48934c4449422e9784eac4f2b11eb4487e

Observation 1d456f56-1889-441b-ba6b-3183b8b6d3bf · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

OmniGen2: Towards Instruction-Aligned Multimodal Generation SAM 2: Segment Anything in Images and Videos

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.774361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:b11fc2ca6205d2959813f40b114c2b0d38999d1429058c6bdaa51422e55ae15d

Observation 91190551-509c-439d-b59b-fd85a7951d14 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation High- resolution image synthesis with latent diffusion models

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.363615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:ac74f6ec939a5434a4fdd3d2b68e1086f4694f9168384116b402f17cf581a1cc

Observation 8b33ffe8-3d3b-4a81-b14e-f01854e1f3e0 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.398728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c2ab053411417bd5292a0788207aef9f31c3b83cf5f8d564b2af92c5a32624ab

Observation 5a859f6b-dc4e-4519-90e8-67c07c6d8e2e · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Laion- 5b: An open large-scale dataset for training next generation image-text models

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.402499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:fa31e6e2a5392e7e66406e7bd5cb5d8649f3ede8b5d393f6bc54213a630bf5aa

Observation 0cb7f001-34d7-467a-91e1-8031047eb52f · outbound

This paper cites Emu edit: Precise image editing via recognition and generation tasks.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Emu edit: Precise image editing via recognition and generation tasks

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.330131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a1993ab1766bb58a3c5fd0d4fe78124512c8b5b307b038bd2e1238bc734c9a56

Observation 301bff5f-135a-4629-bd06-220f1bcd1fea · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T07:52:10.858881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:d0d31a74f7dbf590ed16304fc42489869866221f9d0068029f246a83f9ff6dd4

Observation 8d17d2e5-31b4-407b-b35b-1856094b0138 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Journeydb: A benchmark for generative image understanding

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.382048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:3b45bed694c0a8b979d8a75b7f0f0134bff3a69a47aa02743b15564ae92a6c38

Observation 6d0613b1-39d0-4197-bcba-1f0903ca05ee · outbound

This paper cites Generative multimodal models are in-context learners.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Generative multimodal models are in-context learners

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.421781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:16988aa5257982790d0422f26e1bbe6af0aaeadbd96f2632c02d9c279cfdde4a

Observation 093b651b-a71c-4722-8a66-5191e621d724 · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

OmniGen2: Towards Instruction-Aligned Multimodal Generation OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.949091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:9c980ddb3c4d3fc4c029e79c4aa27d23b8e712819140d13eaa60d255e6db1e1f

Observation f34466e1-4977-4e60-b676-d991eeaa3200 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.951916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:7ffe3358d4a9ab1419dfed931966d85b9011fc98b5eb947bc8769a80c8d0d7e4

Observation 42930488-b5c8-4285-9f6d-7ae72bf657f4 · outbound

This paper cites MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing.

OmniGen2: Towards Instruction-Aligned Multimodal Generation MIGE: Mutually Enhanced Multimodal Instruction-Based Image Generation and Editing

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.958146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5468c5a5c3cd1625f22ff9b9de18ffc36376d6da77e7e13f1c943d1dad3a4238

Observation e5ba5c38-9473-4601-832d-f72cf168f5e3 · outbound

This paper cites Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPO

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.937303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5f126f8bdfa332e4541fcc2ede5ec4892eb2d77b3ee3ff2555e4c1d9ddf8b41a

Observation 77c7b205-9491-40d8-b778-e66da0da3f59 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.945573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:bbc1f89fe091d78a159c100840ea3618f8f28508314df22c7a75254452e2ff8a

Observation c858461c-f18c-48d2-8d98-eae8801f644e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Emu3: Next-Token Prediction is All You Need

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.941892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0e64755d41fa863df2ffc131ac3a26a9e497a72f2a48a27d7e6a6833d7a20582

Observation 140f44c9-653f-4573-926d-dcca532e1a19 · outbound

This paper cites Omniedit: Building image editing generalist models through specialist supervision.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Omniedit: Building image editing generalist models through specialist supervision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.415172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:68a5f26f11641a42e4b825815c8ee8238b31f741d2b4497249a82ad54af6d178

Observation e755530c-fa21-4c72-aba5-6dfdb36e33d4 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.349755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:6ec122ec6f332536b35fcb06a88d98774626f2249bbff04b5ea1b5f3a7c125df

Observation ef199316-fc17-48f4-94e0-96ea07d976d4 · outbound

This paper cites Less-to-More Generalization: Unlocking More Controllability by In-Context Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Less-to-More Generalization: Unlocking More Controllability by In-Context Generation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.924837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c46841b8a446c9fd259b2ae7405e3fe4aef97a97326579efa9b1ec61ed152e87

Observation bc61e83b-2629-4f0f-869e-418c538adfcd · outbound

This paper cites Omnigen: Unified image generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Omnigen: Unified image generation

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.419075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c2807a5bfff9f4074fec8cade8676875d8e5341fcaa7f7b6b1e7d19122e6e81c

Observation 106f55b8-8ca1-4360-b981-428e9da0ba43 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.933723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:dab47325804d1c89256b3ea624cb0894789048f6f953d72d9731bc8b65530da0

Observation b2798405-bc40-48f6-964c-ef41567dfdaf · outbound

This paper cites IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation IP-Adapter: Text Compatible Image Prompt Adapter for Text-to-Image Diffusion Models

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.918010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a1b5e16ed88fde641fe712400e72a93ac2b98996642d2db4d12e8d64a31da5c0

Observation bbaac2e6-aa06-436b-81cb-af82f727f2dd · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

OmniGen2: Towards Instruction-Aligned Multimodal Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.921204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:df90a4441deedfec5d17c570db902f71938e2febcd32f8f9f29af2d8b4c730c9

Observation c70239cd-dae7-452a-ae2f-665b7d5562eb · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Anyedit: Mastering unified high-quality image editing for any idea

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.411807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:2ca6f91c6e9c5a602873b1ba5820b33d4446dfa3e1ef465ce4ef551b3a61e55a

Observation 16e7f1ba-261a-4908-a691-8ec2a6e9b962 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

OmniGen2: Towards Instruction-Aligned Multimodal Generation MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.910090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:52635d9f4d5c90a58c503deb55ddc067310d880b4f91bce9a0437721ed37ac74

Observation c5a54885-27f7-4021-b985-ff7c68482c12 · outbound

This paper cites PromptFix: You Prompt and We Fix the Photo.

OmniGen2: Towards Instruction-Aligned Multimodal Generation PromptFix: You Prompt and We Fix the Photo

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.907210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:229e1b6d09fa595d00c651298b2af2059e3c3cf6e472d6ebe150e2c34e8eefc6

Observation 05ae6f34-1b01-423d-9f86-eb0ea44193e1 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.429054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:9883242b4f9ddaf08d0d67b7b876c279940ef3194d9de0eb3134466a06679ac2

Observation 36226127-c661-4c1c-9528-24bf2e23dab3 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.405950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c53c553fb83a936bf711ba50056c2a6815b77e79367159c933b6b802378a9282

Observation 28d57a9f-f04c-4bf8-87c7-aa11c44bbb4b · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Adding conditional control to text-to-image diffusion models

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.382325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:c459bd788626afb6fa0d98c0e45e70327e2842b6c0400d829aef65efe3dd28b8

Observation 16cd7688-c74e-478d-a161-de57d9098dc7 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

OmniGen2: Towards Instruction-Aligned Multimodal Generation In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.903815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:9f957c9fa0b017a7834230d881e6122d14e681006fb1602b7e8c0a86b9cf36b5

Observation ae88a07b-442c-4b3d-b86c-e65c9e7d154a · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Ultraedit: Instruction-based fine-grained image editing at scale

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.431466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:20fb8eb6cc79595baa3b30023b09ad202da15f65a05967264fbf943ccb9483ad

Observation db41be30-8739-422d-9c6b-05007cc19616 · outbound

This paper cites Uni-controlnet: All-in-one control to text-to-image diffusion models.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Uni-controlnet: All-in-one control to text-to-image diffusion models

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.363668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:a340a2bf40a09a1eaaff03a22f81bc319d9a4dc2064ffd5f95a3c527a00e2c29

Observation dcfbb708-e06d-40af-b6a7-5c90929349a4 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-19T07:52:10.899955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:5baaddbd5163790bf57406034a431046986cd21f4b60a9b2c40d51dd0161bf67

Observation 71813034-7ca3-40fa-bc61-c7a1bcc69681 · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.

OmniGen2: Towards Instruction-Aligned Multimodal Generation Lumina-next: Making lumina-t2x stronger and faster with next-dit

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T08:23:03.378510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:1b69a8c7fdd9afaacab75a82bd59df8ed7d594fbe98e31efc9a6ddfc6052a493

Observation eacc22c5-403a-4eda-8951-f1671aaa3d80 · outbound

This paper cites From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning.

OmniGen2: Towards Instruction-Aligned Multimodal Generation From Reflection to Perfection: Scaling Inference-Time Optimization for Text-to-Image Diffusion Models via Reflection Tuning

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:52:10.896518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T07:47:34.464711Z digest=sha256:0f163713c308d1a672566777bafcb3916d110ac4ea36362dda9fa35f8a471596

Pith citing papers

Observation c67b9732-02e1-478b-8c54-be1a6568717f · inbound

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation cites this paper.

OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:44.468825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:44.468825Z digest=sha256:9e13643461166db32d6b5dc23096ce5c951c9135b3cd6d390cfd104f4bf66056

Observation 57d336e4-0130-43a1-9f55-007a8c728c3d · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 117

Resolution
verified exact
local_arxiv, observed 2026-05-12T18:51:15.772653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:b874a117a2a1e29b65046a204bf2c877ff5e59a6ac0a717ec0c79bbe151db1d2

Observation c6695ebe-3967-4611-900a-7da26f6548a4 · inbound

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation cites this paper.

XVerse: Consistent Multi-Subject Control of Identity and Semantic Attributes via DiT Modulation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:31:39.642465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:31:39.642465Z digest=sha256:ae763efe5de6a682efd9ebc2e0edb9c82546ba697610b6c2449373a559d12680

Observation df0d8873-5fd2-4ce5-9671-a01b95f3a8e0 · inbound

Ovis-U1 Technical Report cites this paper.

Ovis-U1 Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.967384Z digest=sha256:6d77fe735ac1c2c2e33d95d789cf72c94b573233a7ff087a6ca4cd90bb5a9bca

Observation 74e7d532-c840-4003-b8b1-7aa5177ed517 · inbound

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset cites this paper.

GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:05:39.715086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:05:39.715086Z digest=sha256:52017619205561dc6bf9cd2c41d4bfbb2638b660315612e60a8bbbd02bd77553

Observation 0317847f-23ca-44d5-ad3c-a9b2990ae033 · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.247786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.247786Z digest=sha256:d41371fb68820f1c9ffba38fc3f2f9af87f8c9e09d0d2039438b8406f357de0d

Observation 2fcca5d7-f0e3-42cd-aa47-7cb2d58663e5 · inbound

Qwen-Image Technical Report cites this paper.

Qwen-Image Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:29:07.083490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:29:06.883874Z digest=sha256:c1ac1ff408692011f5eec6b9131d6722752d13bdee9820756b195bdf6a8bfb1e

Observation 276f6e35-9182-4dfc-a0df-9f322ea10bf2 · inbound

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation cites this paper.

Echo-4o: Harnessing the Power of GPT-4o Synthetic Images for Improved Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T20:44:16.475472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:44:16.475472Z digest=sha256:0780b89d7611f6ea5042f7094e99860c8c55748064a6506e4cfb616cba4b31a8

Observation bac6495b-4c04-458e-b93f-75e5e925cdec · inbound

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning cites this paper.

USO: Unified Style and Subject-Driven Generation via Disentangled and Reward Learning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:07:24.709716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:07:24.709716Z digest=sha256:1c8ff715cd6043c7706d213e0b2882c74c3c0c2b7742a868e4e69e530b439b1a

Observation 33e764d5-9f95-4abd-a8a4-42e0097ae95d · inbound

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus cites this paper.

FocusDPO: Dynamic Preference Optimization for Multi-Subject Personalized Image Generation via Adaptive Focus OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T12:53:41.973087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:53:41.973087Z digest=sha256:5d3dbc290a3af39522ff21cec7eb4037f371971e21de7d1bf33484cf3c6e0d8e

Observation 0cd68c83-9b9a-4dc1-8953-e73c63c7c53e · inbound

MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement cites this paper.

MOSAIC: Multi-Subject Personalized Generation via Correspondence-Aware Alignment and Disentanglement OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T12:04:28.975261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:04:28.975261Z digest=sha256:7a3dfd352d83252bd9bfc7608978112b8b50ff408362a2ee9da9febb94b53c8a

Observation 5be15ec2-2c39-4bb6-ad8c-3a652241a015 · inbound

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward cites this paper.

UMO: Scaling Multi-Identity Consistency for Image Customization via Matching Reward OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T23:06:09.110105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:06:09.110105Z digest=sha256:fbeaeef3edfff5f82920c23569676b5abf9809c5522a059bae166b8984eba0ce

Observation 84579299-7ced-436f-8ae8-b816cbe7fedd · inbound

Interleaving Reasoning for Better Text-to-Image Generation cites this paper.

Interleaving Reasoning for Better Text-to-Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:44.967227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:44.967227Z digest=sha256:1cab89fdfab0367eb0c6d9ef2217123e48227f2afc2271a63293360d03fff348

Observation b8339fa3-58e9-4268-a468-04193c791a09 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.233784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.233784Z digest=sha256:16cead9a1b0917796a770584b76c655e5bd1d454659b6699962ca7e70f37b701

Observation 41db8f31-f4b4-4a70-b193-2734420b0857 · inbound

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark cites this paper.

FLUX-Reason-6M & PRISM-Bench: A Million-Scale Text-to-Image Reasoning Dataset and Comprehensive Benchmark OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-04T18:48:03.845193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:48:03.845193Z digest=sha256:7c8f44f82909cfac0b685b7d9093c671e0fc7715bda61e388de90b2d0302ec2c

Observation 4604da4f-afdc-492d-8ad8-bf46b607764f · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:40.863193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:40.863193Z digest=sha256:e647a459564cb4f0dd790c197d04723ac4e834f541239cd9fc9bb1c1412a6a56

Observation 68362cee-3233-4960-8091-bf7eb83ddc59 · inbound

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples cites this paper.

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:09.835527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:44:09.835527Z digest=sha256:53653c18796c40b9aa9ee0de73a9aad0949e04c11e88b1fe6d8804d75c20d8f7

Observation d0d8e164-b510-4ecb-9384-2566a2feea1a · inbound

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing cites this paper.

Kontinuous Kontext: Continuous Strength Control for Instruction-based Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T10:47:12.656350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:47:12.656350Z digest=sha256:e8797e7f892f7c51d4e92e75a9c221844e2801d34ceaa665e2e53cdb24e00422

Observation ca619714-e36c-48d0-939b-d8e4c6ae6da9 · inbound

Adversarial Concept Distillation for One-Step Diffusion Personalization cites this paper.

Adversarial Concept Distillation for One-Step Diffusion Personalization OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-18T04:50:53.709691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T04:50:01.364000Z digest=sha256:bb78035cc238d4b6cf969ed32bed03028f1d05839f4ee9607a9f1b413ac071c7

Observation 32282b4f-8d65-4661-b3fc-7d3bb3cb4a24 · inbound

Emu3.5: Native Multimodal Models are World Learners cites this paper.

Emu3.5: Native Multimodal Models are World Learners OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 107

Resolution
verified exact
local_arxiv, observed 2026-05-18T01:12:13.597119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T01:12:13.426640Z digest=sha256:d2362aa0e8e1e9ebe994770009506307a942a2d06e2e421b4df5b2b1f471b031

Observation 9ce26491-9f90-47a9-bab3-09760e8f20b4 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:25.470721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:25.470721Z digest=sha256:5554ca660f860ef63205a5335d5c35e019b357e474bc173cb937984249b10b43

Observation 7491b3a8-d985-448c-a654-e3d6fca166e1 · inbound

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation cites this paper.

DeCo: Frequency-Decoupled Pixel Diffusion for End-to-End Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-17T05:49:08.249750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T05:47:24.669763Z digest=sha256:547838464311400a503e76e26d1dfe3da83d39324df1261181993434cf7e9f6b

Observation fe4599e1-f29a-4df5-9c6c-ec9a0030f03f · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:09.354756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:09.354756Z digest=sha256:e75e25300b12ec21fdad980edf73cc83dfc948a4a4024ebdb36f39b10ac611b6

Observation 11b91c64-a878-487a-bb71-a56fb007dc49 · inbound

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model cites this paper.

AIA: Rethinking Architecture Decoupling Strategy In Unified Multimodal Model OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T04:19:00.546980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T04:17:07.534291Z digest=sha256:38cf4185387a7ac1f0cac6e225fea4d0581122ae5758edf38069307c02d53778

Observation bfe0fc3c-fa9b-4878-8e7e-c1fad62ebae5 · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-05-11T14:08:37.183455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T14:08:36.801359Z digest=sha256:f0d3d32d6d35418b69ee262b977d65879eafe90cd9e27be128a150efc14e182f

Observation 51eaf0c8-952c-4956-a7cf-6ffc95dbe6bf · inbound

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer cites this paper.

Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T19:47:32.893816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:47:32.893816Z digest=sha256:bf81c4cac417c205ed4cf52e4bb2fd7dc0856bf25c53800c305108e8f04ffef6

Observation dfdec00f-77ff-480a-b1b4-e4cb8b5927e5 · inbound

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards cites this paper.

PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency Rewards OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:51:29.402078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-17T03:49:05.489626Z digest=sha256:6d225ec2ceeb1d40cd3a3d41b08e54ff69f51d9aa72ada4ac9cb876a8c023edb

Observation 4132ca51-24ac-4cb3-a6ad-6cece81a4d18 · inbound

Reversible Inversion for Training-Free Exemplar-guided Image Editing cites this paper.

Reversible Inversion for Training-Free Exemplar-guided Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T19:17:16.595440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:17:16.595440Z digest=sha256:dc7e9856e03ade00203455f56c35fd975893bead41ed515afedef11f81c605c8

Observation de38a0ad-99a0-4b45-b251-06a9f1714c97 · inbound

LongCat-Image Technical Report cites this paper.

LongCat-Image Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T08:04:13.062179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T08:04:12.949075Z digest=sha256:fce8d26abec1dcaa04e12c39fbe641fc39327b1ea9fa73f8c4b62489cbf5d323

Observation 9b43ddf8-2a52-4ed1-8c08-51624a0e48d7 · inbound

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation cites this paper.

IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T17:06:57.942013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:06:57.942013Z digest=sha256:13fb67980e1a1b492347ba2bace26915d0e4d54f0d0323af5906393ac1c0e17b

Observation aeec8005-b4ab-453b-a695-d4ce9600be17 · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T22:41:19.264235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:39:32.955779Z digest=sha256:d7b6d2fbe6eacfbd8cf629afb81af100f4a445f6a478b052679bc8df41bab374

Observation ec8b6838-caf7-4b18-a516-544778224030 · inbound

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling cites this paper.

Scone: Bridging Composition and Distinction in Subject-Driven Image Generation via Unified Understanding-Generation Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:37:06.867182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:37:06.867182Z digest=sha256:281d6f231e9923ec6d08c8923d00f72a6f862bd45cf22c79b74c5b2a94256fff

Observation 9fa80018-870b-444a-9a83-36ae5772908a · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.852349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.852349Z digest=sha256:962db2eeb9c5c5ffc75109ee49e3ec6c14f4d61b6600c6926cfca8f681cbd8b9

Observation b3077067-172e-481c-8614-c371964cb528 · inbound

Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models cites this paper.

Benchmarking Attribute Discrimination in Infant-Scale Vision-Language Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.699463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T20:37:18.523951Z digest=sha256:e5c534322a5756d6c93f93e018870e613462961fd56d5761bc5bbe7a5db27314

Observation d46e0128-8174-407f-a488-ffc1ec27817a · inbound

InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation cites this paper.

InstructMoLE: Instruction-Guided Mixture of Low-rank Experts for Multi-Conditional Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:11:11.458899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:10:47.425041Z digest=sha256:8718fdb146624bad1350d084b88dd056ae9a6aad52eee46da85f9e44b7be5968

Observation 892b9c55-ffb1-41ba-950c-15e9530cb652 · inbound

EmoCtrl: Controllable Emotional Image Content Generation cites this paper.

EmoCtrl: Controllable Emotional Image Content Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-16T19:38:20.746947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T19:35:52.111730Z digest=sha256:5b1ddef80be36f4cb49690f673589c760498ad200b51247d3d5e285bdcd76242

Observation 0fde9377-a7af-4fc8-886a-23b5b077b765 · inbound

A Unified and Controllable Framework for Layered Image Generation with Visual Effects cites this paper.

A Unified and Controllable Framework for Layered Image Generation with Visual Effects OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-05-16T11:57:50.137049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:54:46.989748Z digest=sha256:224899ce82cbf1d61f6f072e30e05292833797721d17a45e64be0b39cfb42724

Observation 318707f5-890d-4c51-8954-70b29ae56e3f · inbound

Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation cites this paper.

Hierarchical Concept-to-Appearance Guidance for Multi-Subject Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T05:01:44.685777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T05:01:44.685777Z digest=sha256:2207d5efea74bda89b6b0b2b7472b80535954ab58e44d72cdc248735bcf2ae53

Observation 5f3f31d1-ceda-44ce-be33-d2661a195fb3 · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:10:43.103815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:9ce3fe6d1d305fe0daa1502764950b8947bef8f18f75a2295fcb7f2a5b754afa

Observation cb75bae4-4c56-4e74-afd4-77cfa223ab65 · inbound

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model cites this paper.

DLEBench: Evaluating Small-scale Object Editing Ability for Instruction-based Image Editing Model OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T12:05:04.993012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T12:04:08.487678Z digest=sha256:8e0ed58dea0f38ef5f2b4196c33863cde3f022b4ad54cea17c2f30c51ecbc7e2

Observation a3bdcf0c-3a3a-4055-8e2e-283613f01ae2 · inbound

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing cites this paper.

WeEdit: A Dataset, Benchmark and Glyph-Guided Framework for Text-centric Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T18:25:57.683143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:25:57.683143Z digest=sha256:d6ef55dcdef393cb2177858ac4b63e794740a8d1fbaf0cd9119a6d43e1137b51

Observation a62bb644-a175-40c0-8495-aecb17fb2300 · inbound

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On cites this paper.

VTEdit-Bench: A Comprehensive Benchmark for Multi-Reference Image Editing Models in Virtual Try-On OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T22:40:05.611340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:40:05.611340Z digest=sha256:25b83a3359ff759b542a90aae6885255ef5fc45fa26aa07c8ecb2cf5e94edbfe

Observation 3d07ef0a-c423-43c0-bf56-8ca976ea3dbe · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-13T23:27:11.006580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:27:11.006580Z digest=sha256:111bdef43ee2985beb42cd9744134dc8342c5e420f4e29a2de76396f2a849f56

Observation 9df42cb4-f0c1-4cdf-9ea2-3e0d77a4cf38 · inbound

Demystifying Video Reasoning cites this paper.

Demystifying Video Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T02:33:59.925218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:33:59.925218Z digest=sha256:519e4ac7944be6546857c67ecd696f01d6b229dff05cc8f297fff654148d5098

Observation f2debca2-c864-42a6-b137-d91ef1b07d80 · inbound

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts cites this paper.

FoleyDirector: Fine-Grained Temporal Steering for Video-to-Audio Generation via Structured Scripts OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-15T07:39:50.485071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T07:35:14.257562Z digest=sha256:a5dffc2771c4a5d5fe770236465f93704bf5d7ecd7de18fb3ced232783591e57

Observation 1ba346ef-4557-4dbc-ac8e-74b445b4a9fd · inbound

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking cites this paper.

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-03T02:31:04.995834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:31:04.995834Z digest=sha256:174da5d4302dfe2590dbca53fed806389da4bbcba3375ba58905e56e0e908540

Observation 49d3ce28-6fcd-45cc-a024-537f7821a90e · inbound

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes cites this paper.

HorizonWeaver: Generalizable Multi-Level Semantic Editing for Driving Scenes OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-10T22:40:50.145090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:39:31.357223Z digest=sha256:275371662781a7901755f8e9f2af830a7821465387d627e2a171ff0e1d1db80b

Observation 497b72ab-8d42-40fb-a74e-e3d27832cac5 · inbound

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing cites this paper.

SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:00:50.306536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:23:50.614589Z digest=sha256:c7a7c40664316faf64556aa58c7f2a34b87225edffd43fed27451a744b1c6ea9

Observation 120fb9cc-33fa-4170-98a9-9a6707b7fd14 · inbound

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details cites this paper.

RefineAnything: Multimodal Region-Specific Refinement for Perfect Local Details OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T23:20:54.197919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T19:11:43.172296Z digest=sha256:7ad26fa056e4458da6e09ef3c6361228da8b07c5cf890602236796102d2c9cf1

Observation 97abfa03-6e42-4d1d-ae4c-c3913a7cd11f · inbound

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation cites this paper.

InsEdit: Towards Instruction-based Visual Editing via Data-Efficient Video Diffusion Models Adaptation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-11T06:25:57.652254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:39:59.791758Z digest=sha256:6ebe9c646d84421a40a40ad55224f3b9c94063a5df1457d2e677af62fd3fd311

Observation 2e80d05a-68f4-4e28-8f5f-e15db5bd8618 · inbound

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control cites this paper.

AIM-Bench: Benchmarking and Improving Affective Image Manipulation via Fine-Grained Hierarchical Control OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T10:11:04.299361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:35:49.575589Z digest=sha256:4ecf3d71629252d241b436c859cae078f462331ec93abdfe5f475f094416617d

Observation c25219df-fce1-47f4-a4f6-85975d722c22 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:6428285bec0a9ea52fb3a3ef7212237df797306474d3ceee1743d91a820d54c7

Observation f6778d0b-6094-4bd0-98ef-118084e33d7f · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:21:03.155574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:83a73d89d549682af717e08ef72983df79eacbf879bae44be689a939f23d4cb0

Observation cd658068-bc06-4582-8d91-32ef9fdc7a1b · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:59:55.321131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:eb8c4ae9d01d19d312cf62994bff027427e703a1bc131be1ece515519b51837d

Observation 882b8cdf-cdc2-4754-bab8-c7a7f7572e38 · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T08:50:59.262264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:5c3f284a88701fe4a91aed070fadc813fedba6a6797ef2ea5cb934765cda1bc9

Observation 396e7ff4-29e2-4f81-bbce-b73a27080530 · inbound

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance cites this paper.

FineEdit: Fine-Grained Image Edit with Bounding Box Guidance OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T09:06:00.012081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:15:23.578176Z digest=sha256:db91c05693c91fe7fdd83de03921f293be1da7d4b338139f4488fd49fcb5882e

Observation 3e29a58a-36e7-4657-ac6b-cb29f6e018fd · inbound

Nucleus-Image: Sparse MoE for Image Generation cites this paper.

Nucleus-Image: Sparse MoE for Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:26:00.466256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:30:46.994872Z digest=sha256:fd30c2c2b2351ceb9cf981f8f7247f40c40f96de56e8d528df502af5b7e2e50d

Observation 29eab9d4-3139-40bb-90f7-97473d3f0d70 · inbound

ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding cites this paper.

ASTRA: Enhancing Multi-Subject Generation with Retrieval-Augmented Pose Guidance and Disentangled Position Embedding OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:10:29.281231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T13:49:59.632467Z digest=sha256:57145f125cd640b585547176eb5abb12b5a8aca7a34b845d3a39f280305b865e

Observation 64c77953-466f-41b8-b759-dd3fed8388e5 · inbound

OneHOI: Unifying Human-Object Interaction Generation and Editing cites this paper.

OneHOI: Unifying Human-Object Interaction Generation and Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:25:29.952027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T14:22:30.898080Z digest=sha256:49a3b69b01d22f1587b2c1869ea7e8347745b7404b0c3508590f12b8bab13119

Observation 15f87b3d-ad47-4405-aa46-31afaf05db28 · inbound

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench cites this paper.

From Reactive to Proactive: Assessing the Proactivity of Voice Agents via ProVoice-Bench OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-10T11:45:21.042417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:44:12.373082Z digest=sha256:5806b78628562192fced8f9215d16598e039cbe35bc18a7cd79541015fad2823

Observation 81307ca1-9661-4a25-b07e-3e8dae39c73f · inbound

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs cites this paper.

UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-10T09:08:25.707755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T09:05:10.042162Z digest=sha256:fc5de8eb777f09742511a5fc086c6b6c95fc8e72d2eb04f03061fa94feada50d

Observation bc121b08-a3b3-4fc8-add9-ff73d89f8085 · inbound

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior cites this paper.

DreamShot: Personalized Storyboard Synthesis with Video Diffusion Prior OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-10T07:21:55.030962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:21:45.633310Z digest=sha256:d3cfc356dcd81fb68dc04a32791806a3c1300c855dfd13895294d7b54ae07b30

Observation 0987f470-dcea-4053-bc93-9ada02b2f8f1 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-10T06:06:18.938571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:c041e1d8bf69dd3f4eb1214657f08403024a1f8b2935f80a7758a846a77213b7

Observation b6ea9339-fa88-4e79-b775-9971c7981965 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T07:26:29.650098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:ae01791ca900ac45f719deebdd4541c04f0d9be8e930c658584e7d2e41068033

Observation 25b8428d-e5c7-43d6-a6df-2f329507b3be · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 74

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T16:51:14.205432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-05T16:47:32.853010Z digest=sha256:ca77d3e93dd9bd53e665c6cadb182acfc8926a989fe258b870dd290e33d7313e

Observation d5b66350-6d1b-44ee-9cad-799389f02de2 · inbound

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement cites this paper.

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-10T09:33:41.731697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T05:15:04.561573Z digest=sha256:6efd9702b3a79da32cb3530bd43290e76f180112797f36ade42a8ac357f248e2

Observation 29da7627-6712-49aa-9997-060008403a77 · inbound

HP-Edit: A Human-Preference Post-Training Framework for Image Editing cites this paper.

HP-Edit: A Human-Preference Post-Training Framework for Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-10T03:29:21.957141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:26:15.525307Z digest=sha256:f2cd25e9d368a150cfe52a197cacef9575b7b599f3fbcb1bf81c7dd638f8179d

Observation f412204a-5539-4ddf-a03f-eb4d327c7cea · inbound

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing cites this paper.

SmartPhotoCrafter: Unified Reasoning, Generation and Optimization for Automatic Photographic Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T13:16:03.776146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:10:12.484462Z digest=sha256:b855b4110480096e782d9c07ae2aaaf0ddd54742ed30e51faaa66f6c4a3a46a1

Observation 25659def-5c34-4524-a3e1-4dbaff10dd67 · inbound

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings cites this paper.

MMCORE: MultiModal COnnection with Representation Aligned Latent Embeddings OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-10T03:29:21.663941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:27:50.144706Z digest=sha256:d40341486e0a38a688ad3515b93f5d4fefc0da6b57f163883b4d1f8838fa6f6f

Observation 0f6e9366-bb90-44f4-b33d-c01b29805084 · inbound

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing cites this paper.

Rethinking Where to Edit: Task-Aware Localization for Instruction-Based Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-10T00:39:48.404850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:39:21.872643Z digest=sha256:9327e9dc215fb197de349de81e5344e9298592588eb168aad49abec04f1c68a0

Observation 4bb6c475-8721-4cde-92b0-1ecc64bc29cf · inbound

Exploring Spatial Intelligence from a Generative Perspective cites this paper.

Exploring Spatial Intelligence from a Generative Perspective OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-10T00:49:48.824893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T00:45:46.261005Z digest=sha256:b650749a5191e9285151235de9d93c01bb2c055dcff01c04d06ac6f9a9cfdb2d

Observation b6c6055d-cb62-4252-b193-0b64a05811ef · inbound

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing cites this paper.

Edit Where You Mean: Region-Aware Adapter Injection for Mask-Free Local Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:06:13.737655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T06:47:22.451132Z digest=sha256:d68b4224c443e1f196f9cbbffddab974a99ebd1bb2905fef67922e05fef7fa8d

Observation 36dc96e8-3034-440d-9ee5-ea5454ba3ee7 · inbound

Meta-CoT: Enhancing Granularity and Generalization in Image Editing cites this paper.

Meta-CoT: Enhancing Granularity and Generalization in Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:46:12.247983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:30:28.636915Z digest=sha256:8fb8cc77cd32db454996aa754666374bc51a0b52b9e9b3ffd85c9b28077cf99e

Observation 71618ad8-1a68-4fc2-be81-02083c458482 · inbound

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models cites this paper.

Beyond Accuracy: Benchmarking Cross-Task Consistency in Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-11T21:51:17.895654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T04:08:41.452018Z digest=sha256:f16c68a738878f30da9203fb49cec75ebccb3c1bda3bc440534248a58518a5f7

Observation ee52f3b7-b9ad-4781-99ba-833876aff080 · inbound

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing cites this paper.

DDA-Thinker: Decoupled Dual-Atomic Reinforcement Learning for Reasoning-Driven Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-05-11T23:36:15.095575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T16:44:23.162726Z digest=sha256:58d42db7e524589e66aafc19969ca5376b60a48a09bd15aac01659c6911e8b86

Observation d6b43f2f-f4f1-4bbf-8eb3-638ab79c0261 · inbound

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness cites this paper.

SpatialFusion: Endowing Unified Image Generation with Intrinsic 3D Geometric Awareness OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:46:26.890094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T13:45:53.346402Z digest=sha256:55488cb260dc8ab522ea71b038133e70c71b7ebcbc8fd3e5e4001a72f7f2d54d

Observation b22ca79e-6e97-4824-9f5e-b28743a40476 · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:55:43.682065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T17:57:08.606559Z digest=sha256:ca4e586cc4b0601ef7f6f7f2c961d28fa2b475403104639ff5e8f75103dc3142

Observation ca22cb71-1ee9-484d-ae73-97c5cd52e08f · inbound

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation cites this paper.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.726850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c91ff4d90490cdc0eb63384bc27eb94c77ed71586f69c2cc473ef17476c7ad44

Observation 7d376e86-6a5e-4d49-9e24-016473c1fe71 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 145

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:36:08.010003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:485ad5d8ec535de998c647e4b0a1209a4ffb435a9f07a50cc642196d68141829

Observation eff83de5-2056-4c0d-ad4f-227e41458c57 · inbound

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision cites this paper.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.245483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:938a36e7e697ff88c0dfe691595dcafa878cc7910d5aa928cb7471dc7bd96d64

Observation 71411b96-9953-4996-98b8-e2b3ee92420b · inbound

InsHuman: Towards Natural and Identity-Preserving Human Insertion cites this paper.

InsHuman: Towards Natural and Identity-Preserving Human Insertion OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-11T03:55:55.473389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:08:20.928777Z digest=sha256:fd418ba2a4d2490f32537ddbb69a08e495fa37048681487321b734290d13f42b

Observation 192fc214-3048-4529-bc50-b3b36f17304a · inbound

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement cites this paper.

EditRefiner: A Human-Aligned Agentic Framework for Image Editing Refinement OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T03:45:57.029439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:19:16.677800Z digest=sha256:c6ecf7d07a1f587c54e54b01244321b5e81e8fc07387b7aea532102aa038aed5

Observation c532411a-81bf-42ba-93d7-3a80dde2550b · inbound

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning cites this paper.

ReasonEdit: Towards Interpretable Image Editing Evaluation via Reinforcement Learning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-11T04:15:59.691793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T01:51:42.304051Z digest=sha256:842c79954ff7a52924f6073305aef494ec2866dfa8f013e63b32bedb00385e66

Observation cc490d0a-6be6-42ef-84aa-6f2a880b22d6 · inbound

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation cites this paper.

STARFlow2: Bridging Language Models and Normalizing Flows for Unified Multimodal Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T02:45:57.013592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:44:59.644215Z digest=sha256:9668fa187d6fb392852e4a667afae94147a53c01204dbd0a8b9c57d59ad4435a

Observation 7234789c-dc72-40f9-ad46-ef7ee8317a99 · inbound

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing cites this paper.

MULTITEXTEDIT: Benchmarking Cross-Lingual Degradation in Text-in-Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T08:26:23.423944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-12T01:11:35.480399Z digest=sha256:33cce2bd5abfb835b4ef47c51db06bf11ebeff90ca33568e39a961a8bf9311fa

Observation 74ef5eb8-a994-4b05-96d2-987a0ca67840 · inbound

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria cites this paper.

Auto-Rubric as Reward: From Implicit Preferences to Explicit Multimodal Generative Criteria OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T08:36:24.337092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T00:58:53.626310Z digest=sha256:19d42fedfcbf6e403a55b679e39c68fb92ce1ea8c677a0f22e5899d82bbc29c6

Observation 59881774-674f-4618-88c4-c226a47c88a1 · inbound

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition cites this paper.

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:51:27.757665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:56:43.640874Z digest=sha256:a49641e52901c3d88c66e5a6963ed984696930e6565963e13892d74a552e6840

Observation c729739c-b642-4803-9bbc-15de28edd2e2 · inbound

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition cites this paper.

Fashion130K: An E-commerce Fashion Dataset for Outfit Generation with Unified Multi-modal Condition OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-14T21:58:03.393608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T21:56:10.499014Z digest=sha256:c4d9344103dc5965357dc226e276ad9cbba49cf907226cb53fe37a6314e28e95

Observation b79e1881-614d-4aac-8dda-75c2ab16b2c5 · inbound

Masked Generative Transformer Is What You Need for Image Editing cites this paper.

Masked Generative Transformer Is What You Need for Image Editing OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-12T06:06:26.074159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T04:35:52.355925Z digest=sha256:186f9ba1bb1bf31e4576b201ed39c0f9c16f09ff7b9243a78283c2c43e8fe14d

Observation cdfc8470-3fbd-4dda-8c09-5149de8c4930 · inbound

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer cites this paper.

HiDream-O1-Image: A Natively Unified Image Generative Foundation Model with Pixel-level Unified Transformer OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 54

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:32:29.636734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T07:30:53.939221Z digest=sha256:8263029ebd760a86333212529d26da2dd7c58e5812be6f6cc1d16325cad6ba1e

Observation f4a40299-e364-4a17-8978-75fd63230e32 · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-13T01:47:04.992215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:d51130aea70e08f303db6551544b2c8e98784df25057568c1d01e9bc96a6e4a3

Observation 39636689-405e-4099-bc0f-b71c36f99636 · inbound

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition cites this paper.

RevealLayer: Disentangling Hidden and Visible Layers via Occlusion-Aware Image Decomposition OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-05-13T07:47:32.662951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-13T07:46:50.540528Z digest=sha256:572cedfc806d219ad9864794298259050bcc946cf4dadebe2bdf8c7f78aac55b

Observation 16d0b599-25fd-4eda-a5df-3676c5a9686c · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:02:23.924920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:53:21.851578Z digest=sha256:8ca6e1f96671e2897ac264696ef13b887a32cbe66774a4fbf37e91ae926e5b87

Observation f193d64a-74da-4eac-ae58-191a2e9da181 · inbound

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation cites this paper.

UniCustom: Unified Visual Conditioning for Multi-Reference Image Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-14T22:03:03.399513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T22:00:01.349754Z digest=sha256:e01766bdac638561313c3851d09c7021e9d6188657f0ed0bdd817ff645cd20a8

Observation 0b1bfc16-6493-483a-810d-131246f4bee9 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-13T06:07:22.929062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:02:39.114660Z digest=sha256:fbfd174edcff9030429a32e9c546a8ead9f6109329816a0eb2d71674ebbf78f1

Observation e9c6ffa5-2c1e-4d62-8035-7e5b067cd766 · inbound

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm cites this paper.

Beyond Text Prompts: Visual-to-Visual Generation as A Unified Paradigm OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.583707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T22:18:55.933311Z digest=sha256:9cda10d3d5c582fad2ab33cb9acd5925a15be5860428c39fbc0e1b58a8c77ec5

Observation 4f78e979-c64b-43cb-990a-aa9b56877b56 · inbound

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation cites this paper.

Images in Sentences: Scaling Interleaved Instructions for Unified Visual Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:52:22.735182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:48:04.997796Z digest=sha256:d920902407a158c3ff2d06447f5ef40f7268647f359bfa9fcd4efe07a4847cb2

Observation 80bb53d2-99c2-447a-a49d-f0a9bd9b621e · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 141

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.551051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:bae269239421521ae0c48b40a777326e91ff674a7e32ccd7cf3379a4c6f47274

Observation e01e0ef2-ea63-403a-82d0-c19828723b73 · inbound

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling cites this paper.

Edit-Compass & EditReward-Compass: A Unified Benchmark for Image Editing and Reward Modeling OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-14T20:22:54.661439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T20:22:20.464966Z digest=sha256:347e1562dc07e16edf12d6e675e6f19ce75a89c6b217366032e660ab0da46041

Observation 4b2c1e66-1d52-489a-9c4b-96e5a9e491a3 · inbound

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation cites this paper.

Early Semantic Grounding in Image Editing Models for Zero-Shot Referring Image Segmentation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-05-14T19:47:53.624237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-14T19:44:36.511647Z digest=sha256:356890eae939ca6ae207ec802fd4a6f43fd4ebfa5d3ab1612fd35f1553bd3f61