Pith. sign in

Paper Citation Record · LEDGER

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

As of 4 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2605.05781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.05781 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T14:48:22.805268Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T18:56:58.714370Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact39
  • verified fuzzy29
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 10bb8fb6-edeb-4e17-92d7-3e70b39f4d85 · outbound

This paper cites GPT-4 Technical Report.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.180724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:3dbf4228f1688edcfdfd96b3c899ba3bfb3636c12cda50b9575e6e334cacae31

Observation 5b152c83-e36c-416a-8a01-481ce1336fec · outbound

This paper cites Qwen2.5-VL Technical Report.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.317678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:1cb26ee3d183db39a7dfc035a8f9c22c0f32b6dbb70561556a4a7f72282179f9

Observation 6bb45bf8-0403-4638-891c-dae03592dc85 · outbound

This paper cites Black forest labs; frontier ai lab.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Black forest labs; frontier ai lab

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.309576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d60f1e88dbd8b3695aeaef40e2919d1610d71d921c87fcc0646dc34430e3fee6

Observation 3eb67d1d-55e6-4c14-a0bf-5f36e1b0df3e · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Instructpix2pix: Learning to follow image editing instructions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.243160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:cd778cd5b584f1e5d501905b06769924dce6a8281352c8f45d6ba9753bcfe8d8

Observation 55547654-a559-4f93-8355-ed0627988905 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.302599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:c59852c132545fc546a0b3d9945f9e3736d10a2586fb22258d087937600b7898

Observation dd41f6e7-56e7-4d1a-9d75-6a982f0a017b · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.290503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d62b9cd42184c8faa6c32d0722bcf761fb9ea4fe6b276a728c640ed4068ffd19

Observation b031ab72-d0c4-4ddd-93a2-57750730bffb · outbound

This paper cites Thinking with Generated Images.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Thinking with Generated Images

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.130970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:01254de290f49324920dbe2c45d46219cf54729c4ef924ae755fe96dd2c833ca

Observation 345cb79b-2f2b-487b-b7da-6d8d70ce0dda · outbound

This paper cites EditMGT: Unleashing potentials of masked generative transformers in image editing.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision EditMGT: Unleashing potentials of masked generative transformers in image editing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.223010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:a9568375c68d6edade615d4de511048339ba412273ddf551a9bb990df1503d0a

Observation ac2b1d96-5cd7-4cc1-856d-a176cdc78958 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Emerging Properties in Unified Multimodal Pretraining

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.205130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:298fb49eba25ccc875425de08244269686e51b6151e11e4fde4bfb495503185e

Observation 25f4e5b8-dd67-4e20-b310-0009e062f9e7 · outbound

This paper cites Dreamllm: Synergistic multimodal comprehension and creation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Dreamllm: Synergistic multimodal comprehension and creation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.298640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:228d51466afc5e363b35f4de2abbdfdbe27285f73f5054d704ae23aab9c4a4b2

Observation 7c6f48bd-bc1c-47b0-af80-664d9878c89c · outbound

This paper cites Scaling rectified flow trans- formers for high-resolution image synthesis.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Scaling rectified flow trans- formers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.302562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:a77a0cd680aad712ccc0b1e9092dc8ec809438d29e29d41b0eee3661d660b9b8

Observation fe2be6f1-c4be-4b94-bb9e-1d382f0a71f6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multi- modal large language models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Mme: A comprehensive evaluation benchmark for multi- modal large language models

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.253527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:66346bd9663a4a01601d6b90be51006a912c82f2add3459504d145bf282a0b97

Observation 1d2bbd97-d671-4132-98c7-b617d66971a7 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.519178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:55d1be657d147be41060381073fae538b60d1154c3aee1cd267ce3092ee7d614

Observation 4722ece3-867a-4179-ad36-e03911eace4e · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36:52132–52152

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.281241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:dd2da0418ca93b46fed8182a5d0ca04fd93fb808e4fe97529a4bb65dd12602f5

Observation db8c475d-c299-4cbd-8777-8a8342d33458 · outbound

This paper cites Vq-va world: Towards high-quality visual question-visual answering.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Vq-va world: Towards high-quality visual question-visual answering

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.193962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:a36e1b5b603e1fcb5c5e348b33797f7a1a6db22c12c7dc7e4c233409475dc8d0

Observation 3395922e-dcd6-44ce-b10c-7201a42e719a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.158443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:2124cea215e8e1e18b30b933b4ada5b038c0a72b6fd255c778fa72166e163ee7

Observation 835d9f7c-2081-4d8e-8261-9ccc71bd7f2a · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.312366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d27fb3933b949e15f90ef49d8a1f443e9a3c3903f1c6b52c50d5ae45035ddc1b

Observation b5a41f77-f266-4959-815e-7a08135887f0 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:43:03.755490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:40b1289ce2451273ac57914462aeb00676e98a1a33c7885d3e8daa14ba01abac

Observation 82b1dd93-f2a6-43dd-bf4b-7decbdbfd8af · outbound

This paper cites Anyedit: Edit any knowledge encoded in language models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Anyedit: Edit any knowledge encoded in language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.259126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:061c4a1ffd4ac67514fc6c4ec7934d9ab2736bb5d08dc8df68317f371618b162

Observation 6445a212-8b63-4809-8a5b-c1e80775daf9 · outbound

This paper cites yes" or.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision yes" or

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T18:41:10.240465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d1c01a57f6f17de4558ebbb1f3130bddd703642546456319c011866303781898

Observation 9cbdd6be-e091-44c7-be74-ad6263659532 · outbound

This paper cites Eq-vae: Equivariance regularized latent space for improved generative image modeling.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Eq-vae: Equivariance regularized latent space for improved generative image modeling

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.249165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:1ecc721ac31c0d92a6a0f1ffe8e5d938755cbc684cb5b25baeec84271b208889

Observation e1138c00-2eb0-4282-818d-b3b7af650329 · outbound

This paper cites Nohumansrequired: Autonomous high-quality image editing triplet mining.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Nohumansrequired: Autonomous high-quality image editing triplet mining

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.336229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:c13f6d332dc2fe4a168a9911dbbebf1dd2a32f17a9b92fafccf7be20f792ea01

Observation 6b577421-5d82-4740-a92e-a7b9aefadf41 · outbound

This paper cites an unresolved cited work.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-26T11:37:37.315049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:7371b33fde2c81276c519e70e1515c23ccc694490ad3b3353372f25ac115461b

Observation dff5991a-b75a-49ce-a82f-60ffae2bc574 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.413644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:11ddab22e622d4261b323dfc3a6712440bf8dbab991f98dcdb7ae492ef34a1e5

Observation 4bbdc33e-89b1-425c-b4fb-26ad23fccd9d · outbound

This paper cites Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Repa-e: Unlocking vae for end-to-end tuning with latent diffusion transformers

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.434289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:54a33d7d94eb99b5c684dc27fa99feb59caa6e7044338214c6e41aa39b5b647c

Observation af554a6c-01ac-4b7c-be23-96ac06311ceb · outbound

This paper cites Imagine while reasoning in space: Multimodal visualization-of-thought.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Imagine while reasoning in space: Multimodal visualization-of-thought

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.324124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d392700a6201793fbef44468257b4e45dab17debc41ac2bd56baaa4175850770

Observation 6c8f061d-d5b0-4af2-8f57-cf1830348974 · outbound

This paper cites Onecat: Decoder-only auto-regressive model for unified understanding and generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Onecat: Decoder-only auto-regressive model for unified understanding and generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.310857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:0860649284724a290d8b5bfafb685f8e4f05c8f1b13c602b3ba5934783d4e077

Observation d851a8fd-8365-4267-9b69-c12289c8e7a4 · outbound

This paper cites Autoregressive image generation without vector quantization.Advances in Neural Information Processing Systems, 37:56424–56445.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Autoregressive image generation without vector quantization.Advances in Neural Information Processing Systems, 37:56424–56445

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.278250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:feba1e66740fe4672c98fe2bada0bbc68c9f452fbc7a4413d19beaab30a93b10

Observation 68ca3073-0111-4815-808a-da44b132fe78 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:45.158778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:17bb25f14312842bef8fe19e5f98d333d6346e0c7b1eb54a257f55da5caf1e0d

Observation 1ce28159-7216-4d13-b655-64a974fce6ee · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:883078c383d9edc4b97b95c4ac9903fdbd3e66d09f9a1c105385703d13c0c6c5

Observation 8af7907f-c4a4-40a8-a6bf-dac4afd7dcad · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T17:34:27.304691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:ac92eb6b02544a9db5d752efc7b008847a56c5355e30796266f1f28456811805

Observation 9f026fb8-6fe4-44af-8459-a06fc3ba381d · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.236742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:10204c6eaebf115310307f12bdfba42647bdf903787c7bfd522cc2fdb9dfeaed

Observation fb46326a-f9fc-45bf-a142-7a6622433d1b · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.232334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:5cade7624e948920c284484dfb2ec84c9b504c1423830d5d4fbfe2fb831ec4bf

Observation b29695d1-52a0-4c5f-a2fd-44d032af015e · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.262177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:183f9bbcee9f29bfea892160eb9bf3c8ace8c95d5e74785c68126d06578a5dc2

Observation 56ff0cf1-1875-4963-a979-70fe6d2aa5f9 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:f5ab2892560acabb6e465b1d783b11b64680bd696e3148f0f88aa7a9799c6f5d

Observation 492d389a-2974-4944-a700-9c47885e1172 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Transfer between Modalities with MetaQueries

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:49:23.311693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:f90216784a73c5c3f8b8f298724465afaf8d0ab5b274faf2b26fb4a783bb1b14

Observation 87d28391-0fc1-40c1-b029-ab73b9eaf232 · outbound

This paper cites Scalable diffusion models with transformers.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Scalable diffusion models with transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.233524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:3fcaf7cab5e09a23887048246e64ff75c439408e15130f43fe5ae4da1ea7072b

Observation 14407db3-da7a-4ce4-a730-124670d26dc8 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.327112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:a8acbb1c448fee847b8422bca3fd812724186d201fce79280710128fe8bfde75

Observation bde48f28-98ed-4f07-8e5f-03daad8248ee · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision High- resolution image synthesis with latent diffusion models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.240403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:2c0c31a8c536e5716275d76bf59607a096339f0b979d4fdb01b64588bf6b3e34

Observation 2adbbfde-a18c-4541-a393-829bbbfa22a8 · outbound

This paper cites Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Laion- 5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.306417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:518e8b7c777b97c5c6304145ee251be164c31d6d05e3fdae02e9046d671ce0b7

Observation 59687e1d-c7e5-448e-aaab-aa817c4483af · outbound

This paper cites SVG- T2I: Scaling up text-to-image latent diffusion model without variational autoencoder.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision SVG- T2I: Scaling up text-to-image latent diffusion model without variational autoencoder

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.409875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:70ae82960f78c8092e9a1758aee3e3dce72ce41cd7c8169b23f60381e266626e

Observation 3e4e0ffd-964b-45d5-98c1-968bc43bf482 · outbound

This paper cites Latent diffusion model without variational autoencoder.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Latent diffusion model without variational autoencoder

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.373192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:cf9b7d279201e2f534b86595c57f452ed13b8dcec3179aef218c591634bbb1ef

Observation 463a70e8-fdc5-489d-9f5f-e494804be603 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Journeydb: A benchmark for generative image understanding.Advances in neural information processing systems, 36:49659–49678

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.271897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:e8decdcf1197a94acf546410f3d10363295254c01a395e617c4062544da74444

Observation 8f089a44-edc8-48fc-9eea-36595f83fefa · outbound

This paper cites Generative multimodal models are in-context learners.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Generative multimodal models are in-context learners

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.246165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:14f38b18188bb5cc1ce8cfd17e4a0c7abb6b06efbb4a5b7f56aa5533adb48424

Observation 57bc9132-c8bc-4d6a-b52b-d6bed248d7ab · outbound

This paper cites Exploring the deep fusion of large language models and diffusion transformers for text-to-image synthesis.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Exploring the deep fusion of large language models and diffusion transformers for text-to-image synthesis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.321254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:0434a2268b84e57f9af2770a00795c03132d946d2bca90f27f5168e40e74e2d4

Observation c11edd16-8839-4ce7-89d2-5ff81514a8c4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.383062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:e80e10170784b61f8505f3530f6ea70b7e6690f082fe569b89ca2805d6360527

Observation 5cce2100-9b19-4717-9a7d-20fa9edbd5f3 · outbound

This paper cites NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision NextStep-1: Toward Autoregressive Image Generation with Continuous Tokens at Scale

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.111591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:8b6d5c56795da17d028e007a1fa043a420fb77b91a4e577522ce93cd0a8275d9

Observation c47e8962-5214-41ae-9763-ef8ba6967329 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:00804fb23dbd54b19f20ebc57e44bd618347ea05ce7eb44b85896c616fc6d8e0

Observation a1b3da3c-d7d8-4e20-8b2c-98a41ef3bcab · outbound

This paper cites Scaling text-to-image diffusion transformers with representation autoencoders.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Scaling text-to-image diffusion transformers with representation autoencoders

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.199620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:4f0903ac7cce919491838b3455c69339f141e390807f0d937692e430ae4f82f8

Observation 76fc4755-77f5-45b3-840c-3c16c7a48942 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision LLaMA: Open and Efficient Foundation Language Models

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.117243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:35743e392d4d6ee30c3102cbd1431f6fdfd67d630559a8e620d3c584c7a63f52

Observation 6daf6267-1aa6-4fa7-9662-65bfa5761ca7 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.123584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:cf94437a0e08c615f7de0726292ec293b49cf99f4bf5a9382fb135469dd290f7

Observation 4a52d568-e384-4347-a93f-0a26a61eafd5 · outbound

This paper cites Reconstructive visual instruction tuning.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Reconstructive visual instruction tuning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.318130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:c6fc04320570d31b5de54e53e2b24eaee370ef3cedd3b701977869358459b38e

Observation 3a9ca61c-c587-4eac-b716-b84075c2b862 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.107136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:d72ac0aae6d5ad05d62c4218ef99b86155afcbe2c51056480189def1aedee69a

Observation 0c810bcd-36dd-4c19-acec-0f22177b4540 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Emu3: Next-Token Prediction is All You Need

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.260713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:1d701c6385b418fa549bba43ede7d6e61558fb9466692d62ded407cfd3e3d856

Observation dee93dbd-b567-4bab-9bdd-54cfba94765e · outbound

This paper cites Unigenbench++: A unified semantic evaluation benchmark for text-to-image generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Unigenbench++: A unified semantic evaluation benchmark for text-to-image generation

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.421059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:1da3f0978237ec4a339d9626ba52e82852cc32dbc80d5b72fe3a881976f903ea

Observation 6ea1ef48-e134-4a2b-852f-f60129561aa1 · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.268462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:09d4577dff32824ebaae1fcf1e3e6380a8f2732104ac0b16a344282e5459bdc8

Observation eff83de5-2056-4c0d-ad4f-227e41458c57 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.245483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:57c892f31b0d0f362d5028b7c938bc050755777f64c80e9e09930216ee085ae0

Observation 467a9fa6-95b4-4b21-9f6e-c43f35fdd2ca · outbound

This paper cites Visual generation unlocks human-like reasoning through multimodal world models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Visual generation unlocks human-like reasoning through multimodal world models

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.363848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:f8edb78466aeda596e0d3798cd68348bbbcec28766889ba0c87cda469796f1e1

Observation 07e9c6af-4246-4a18-bc7d-2f2ffbb24926 · outbound

This paper cites Liquid: Language models are scalable and unified multi-modal generators.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Liquid: Language models are scalable and unified multi-modal generators

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.292143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:c22f8d403855226c622a4b52ac7726cb6ba39db7c27c67c48b36500ad9e80400

Observation 333d691b-d2e0-4bd4-9cd5-d6ba23abcb42 · outbound

This paper cites OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.358702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:bfa5f9d5929cdc4728085f03092735f5357a1bc2a7b30545302244aa29a530d1

Observation 56957309-bc53-47df-b6f0-59e2b949f452 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:26:21.484895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:b751e7a156fb48b8dde375291834d0a943bd2de59405f60b88110c48d2975694

Observation 7ce7be28-b763-4aa1-bc5c-c7508364adf8 · outbound

This paper cites Kris-bench: Benchmarking next-level intelligent image editing models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Kris-bench: Benchmarking next-level intelligent image editing models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.288823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:fdb76924efcb045efb2a7ced4a4556d54dc3b77d18301ae8206f6baf6cd2455e

Observation 7e57cca6-0d17-4dc9-8172-6ed1e736f48e · outbound

This paper cites Omnigen: Unified image generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Omnigen: Unified image generation

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.256416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:6406e3bc0530055486131b0094fc54c2efa5868117c07c2b3c243d8ca7806799

Observation 7af36f02-f6a2-47fd-a249-62fed1bf038c · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Reconstruction Alignment Improves Unified Multimodal Models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T02:15:36.738364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:135902f9177d34af72177ab87b5c124475acb057093295df2af1f94a3aaca0f3

Observation 11931339-98f2-4a2f-81ed-6975f2a12951 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Show-o: One single transformer to unify multimodal understanding and generation

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.295419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:1c8ab25c4ab4332a00b320e3a9f597e86349b45df678f2e43c153456d52937ff

Observation f198310f-7118-497e-988c-b0a584cf4aa8 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Show-o2: Improved Native Unified Multimodal Models

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:16.424084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:5e2a4532b910e56fce410703e855f0a24bbf8f1723fbe58d48a09a1f3d7f92eb

Observation 88694500-2b08-4d55-9292-8d2b4b8d8545 · outbound

This paper cites Qwen3 Technical Report.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Qwen3 Technical Report

Reference 67

Resolution
verified exact
local_arxiv, observed 2026-05-11T18:41:10.325345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:7858dfa4482db9673291eb9f3aa868a34ba75e9e3b00996f64133899efbe5d97

Observation d49e99d1-c05f-4a1d-bfec-ce135a40c3be · outbound

This paper cites Imgedit: A unified image editing dataset and benchmark.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Imgedit: A unified image editing dataset and benchmark

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.265623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:18285bb6829584c89fd9018df3cc3365aca04ef03ece77b7b0a7e86ba5afb19f

Observation 2814af93-2e61-4341-be89-ab3f56295e66 · outbound

This paper cites Representation alignment for generation: Training diffusion transformers is easier than you think.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Representation alignment for generation: Training diffusion transformers is easier than you think

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.285009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:c9552c7a723609e7801c3e379107bde1ba492162745adab9badbba3e4428c3d8

Observation edd8b983-aa47-4415-a898-928e5a832812 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Magicbrush: A manually annotated dataset for instruction-guided image editing.Advances in Neural Information Processing Systems, 36:31428–31449

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T11:37:37.275381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:e0e31fa663bb987a159d6453554c2b455bae3c8ea2c033e91fcd29ead817b501

Observation ea7ae9cb-6703-4f27-946a-d11d66c078b0 · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision Diffusion Transformers with Representation Autoencoders

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:34:18.071377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:e208901c010d13a74c45e3619ef6fa10b7a1bca4166ecb63f6ea1bebf860eccd

Pith citing papers

Observation 34c57d94-5c67-40a6-93a0-01c71264e2b1 · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs Steering Visual Generation in Unified Multimodal Models with Understanding Supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:58.714370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:58.714370Z digest=sha256:3eb02e9ad620e6856a3aa0b957432125bf819ee014dbabc6798c8a0df4e51eee