Pith. sign in

Paper Citation Record · LEDGER

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

As of 11 August 2026, this Paper Citation Record lists 100 of 104 outbound references and 3 inbound Pith citation observations for arXiv:2604.28185.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.28185 v1

Coverage vector

measured 100 of 104 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-07T06:38:04.459129Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T22:58:38.610609Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T23:04:01.587743Z

Reference resolution

100 of 104 outbound references displayed

  • verified exact78
  • verified fuzzy5
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch17

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0566a21-f8cf-4eff-9d9e-a1eba39f51b7 · outbound

This paper cites World Simulation with Video Foundation Models for Physical AI.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling World Simulation with Video Foundation Models for Physical AI

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-12T23:01:14.323574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:931b8ff15d35a954121390a4b520343228cff55573195fac68a442d58d106584

Observation cc7f424c-9e3b-4239-b2ad-4fefa91f2792 · outbound

This paper cites Wasserstein GAN.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Wasserstein GAN

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.809726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:5978398225c7f491c8c334b39b915c519dbd31814dd6b9fb222c72ca788f8107

Observation cb126e53-c21e-4bb5-bac8-209655f460fd · outbound

This paper cites Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Dream to Manipulate: Compositional World Models Empowering Robot Imitation Learning with Imagination

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.818836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:3dd2434395fe1856516cd4b5fd6a9e66e9043c3a5b3ce607d5abd16a72ce8d6b

Observation 161e0aa1-6ffe-4fc3-bd71-aa48e77dc7ca · outbound

This paper cites DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DRAGON: A Large-Scale Dataset of Realistic Images Generated by Diffusion Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.802249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:02f43532bc362b7c2405fdf5705afa0b45efdbc171839a42d43e3241999de477

Observation b6b38410-aa49-4313-914a-1f1061c7399a · outbound

This paper cites Training Diffusion Models with Reinforcement Learning.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Training Diffusion Models with Reinforcement Learning

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.795740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:e22a29d72200e04b6ae6f5d38cdd0f78400ada675fa91872bcad7dc3c183e1ed

Observation 5c925bd2-ab32-417e-810d-128f4e5a6e8e · outbound

This paper cites Large Scale GAN Training for High Fidelity Natural Image Synthesis.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Large Scale GAN Training for High Fidelity Natural Image Synthesis

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-13T12:11:19.483934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:49006e4e5e7d4bef61f7eccc7dba9292f312f9b2ad0e876f50827cc6a64972f9

Observation 447a3ff7-4aba-47f8-90f5-f57ee42ad371 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.778015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:665febc76cbb9d11afda667e4e87e21f1916911b4087ba0bbdf6d23f07d094da

Observation b5052f5b-c279-48ca-b000-4733fc5e2f83 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling HunyuanImage 3.0 Technical Report

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:32.945705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:8382007683bf270b3f1d689384f7c49f3080b1782601ff5db28052d5a03dfbd5

Observation 167f1811-6de5-408b-9f63-8d2961f1a00c · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:16:29.022871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:dfe4da23f2ab5f26554cdcd5cd7f4a54aa5e1e67270bba2891f71e084bc31574

Observation ea242fe6-5103-4b7e-a0e9-b83d32f2801b · outbound

This paper cites Large Video Planner Enables Generalizable Robot Control.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Large Video Planner Enables Generalizable Robot Control

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:29.014594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:67bdc61a22b2624466721998d325cb167936363e5e8a061556f26174a79b85ff

Observation 3cd79797-8de5-476e-bfad-8999bff9af34 · outbound

This paper cites δ-dit: A training-free acceleration method tailored for diffusion transformers.arXiv preprint arXiv:2406.01125, 2024b.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling δ-dit: A training-free acceleration method tailored for diffusion transformers.arXiv preprint arXiv:2406.01125, 2024b

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:55:12.109816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:a7e64a696d81a6514fb11f6763d1027469e8eab2e092ce7bc02280dd0427d8b0

Observation d2322296-1494-42fb-960d-80e5f747827d · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:21:27.672546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:a43bbfed10898041424779711045482210ce3bfea3b225c32621e272e13feb4a

Observation 5fb7c4dc-b0ba-4d2f-8943-97f00f22349a · outbound

This paper cites PaddleOCR 3.0 Technical Report.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling PaddleOCR 3.0 Technical Report

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:24:28.568269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:a6b73051d40678513c8954c1e55fc84e19bb13c34806c75ced2a12dda414cc16

Observation f5fa5d4e-e8d2-4e73-a450-ae8f9f27fcd8 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Emerging Properties in Unified Multimodal Pretraining

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:21:27.655502Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:d6351da9aa1b174353c42752bbb070c509aabfceea654f5c3660f264820f1f29

Observation 30ddfd5e-d068-4b86-b19c-cbe57bdf80f4 · outbound

This paper cites PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling PRISM: A Unified Framework for Photorealistic Reconstruction and Intrinsic Scene Modeling

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.952930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:66c61d274f8d9761910bc90fb1e38bb0181c5bdebc730c384bae29c446c458a1

Observation 421e7f64-949a-4261-9e43-ec4c28cd530e · outbound

This paper cites and Barry Zhang.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling and Barry Zhang

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T09:59:02.402646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:26a50acc875d0be22d82ca239bc1f1ac880b009160e7fdc995aba4db498a82ff

Observation 1999ac07-cae5-48a1-add8-8893eaf3d8e8 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:21:27.825596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:6f8358b60facc810f84c45643aeefc8e2fc308a3b273bd8ed76f93a446f69017

Observation e88ef3af-694d-4d98-8b49-aad4eda53f95 · outbound

This paper cites Tinyfusion: Diffusion transformers learned shallow.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Tinyfusion: Diffusion transformers learned shallow

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.934530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:0ef6a6f7e0dc6c9ef214bcad24896f03ac382d20139711d0991710a3092b50a7

Observation ada0c6c4-2647-44ae-8914-4add60e89585 · outbound

This paper cites One Step Diffusion via Shortcut Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling One Step Diffusion via Shortcut Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:40:55.602339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:a0d4c92a43b7160f3a80c2f3a1c0b985c1fe4444991ad96a10c8a81c3a2ebafb

Observation 4271ccc4-29ce-4612-b990-e279b1c9f807 · outbound

This paper cites arXiv preprint arXiv:2506.01943 (2025).

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling arXiv preprint arXiv:2506.01943 (2025)

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.924869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:8549d284c6a12c9bfc5e9e1d0c3ba2d54cfc1543343abbb8d22fafd467e0e154

Observation 9d898da6-e13c-491c-b0bc-b38ef5562ee7 · outbound

This paper cites An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling An Image is Worth One Word: Personalizing Text-to-Image Generation using Textual Inversion

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.863493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:9f60c1e02e1a14f6182e254ce9556903f64b00f31aa46e5648619721d2b8c03a

Observation d3d9052e-f055-4ade-88a9-9eaed7c81278 · outbound

This paper cites Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, et al.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Yu Gao, Lixue Gong, Qiushan Guo, Xiaoxia Hou, Zhichao Lai, Fanshi Li, Liang Li, Xiaochen Lian, Chao Liao, Liyang Liu, et al

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T04:55:12.113310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:6412878eb5d911c6eac602a4beccb696ee102a5cc288f78c70ba2e4ee44d06b1

Observation 2c7c1a74-6b9d-48f5-b189-7483a52f626a · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.748056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:78085cb6ab711f6a7e9f312f76f277b2afbb92dfa1c210cee9fb0994c0af1d60

Observation 707fa088-1427-4dcd-8293-4014775e82f0 · outbound

This paper cites Mean Flows for One-step Generative Modeling.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Mean Flows for One-step Generative Modeling

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.762653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:924d8dbc680f77d2e76710d220698e7b2d49f191511dcc11536ade4ad1f85afb

Observation 2ef460f8-c199-4d91-9f88-13e9fdb47be8 · outbound

This paper cites Generative Adversarial Networks.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Generative Adversarial Networks

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:04:41.117563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:417a647ac3928a0acbd2b916f23f3b214627eb503bf85ad456e8d5a5ea504e65

Observation 8c3754db-dd38-441f-b252-095f165ad71f · outbound

This paper cites Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.986490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:0d1d5f26d7eeb8db4b09f1f52ec973bac1d24781ab674601107ee4f665bc578f

Observation 81cb7cea-0d45-46cc-bbaf-3efe77f69408 · outbound

This paper cites Gems: Agent-native multimodal generation with memory and skills.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Gems: Agent-native multimodal generation with memory and skills

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.805706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:d7ae1ca43fbcb0dcdc16cf2e289c90a4d0bc287f878f1352b14856a73f4780d2

Observation c77df711-f5dc-40ec-b309-f405ce53b9a5 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Classifier-Free Diffusion Guidance

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.968023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:2211de190e4b3524bf9d2f6c01a23de3879901e9ad454ab516804d52906c8f4c

Observation 7d324dc8-9184-42cd-9293-d267371080fb · outbound

This paper cites GAIA-1: A Generative World Model for Autonomous Driving.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling GAIA-1: A Generative World Model for Autonomous Driving

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.995031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:67c33d55a43a2e7df8ae2f1b601e43149d59be37829200d1e452402e85aed36f

Observation acb75062-90bc-48cf-9c0b-d8abe9ca9d49 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.686377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:82ecb850d46ef977ec186e898b698c04eee096e51cac35c89b010b77d4b1ef7b

Observation 7308c216-8f56-4c5a-bf1c-e260093ea199 · outbound

This paper cites The GAN is dead; long live the GAN! A Modern GAN Baseline.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling The GAN is dead; long live the GAN! A Modern GAN Baseline

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:28.948203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:77891d40edcf7bf1f98ed2807b9c0a690c41340540a14c50a2e2d4fd4e1961db

Observation e04d23cc-0d2d-45a0-a22e-d9cdfb33cdfc · outbound

This paper cites DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:07:22.650116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:200cd2df8d49eb78b34b69bf01a87d98f091a15a5683b438f5697eced222f809

Observation 9f7719dc-460b-43e8-8dcb-4178c808cea5 · outbound

This paper cites DreamGen: Unlocking Generalization in Robot Learning through Video World Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DreamGen: Unlocking Generalization in Robot Learning through Video World Models

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T23:50:45.800685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:3696ecbb18dbe8bdf6cf3c2bf343817be155e9e0a0528e94bde26dd63aae8c17

Observation b2bdeb14-5dd8-4543-a03b-9150eb3a09a5 · outbound

This paper cites COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling COLE: A Hierarchical Generation Framework for Multi-Layered and Editable Graphic Design

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.911678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:7a2db0a9cf0a8226d033e951d90caff5fc9258e0ba266482215e24bb32fe8cbf

Observation 4bbe8f97-9cc2-4d92-80ad-6747eb1045df · outbound

This paper cites Lego-edit: A general image editing framework with model-level bricks and mllm builder.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Lego-edit: A general image editing framework with model-level bricks and mllm builder

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.867149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:1f9f20556f326cc4febc60a06de7afefcc991ada1892fe3fd3922a5db89db967

Observation 39a1bbd0-4b9e-477f-8d14-e2baa6355eaa · outbound

This paper cites pdf, accessed 2026-04-22.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling pdf, accessed 2026-04-22

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T09:59:02.396163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4c50b798801cba6bbec40c6ea204c0f2a646c1f040102c33b334aaa66c2b231a

Observation d402a5f3-ecc8-4b99-b97e-389512a1a33d · outbound

This paper cites A Style-Based Generator Architecture for Generative Adversarial Networks.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling A Style-Based Generator Architecture for Generative Adversarial Networks

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.890876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:d6fc683690b2e17b2e54c965a1ac1ccdfd2ac921b8e2358c7bb8cb7bde2c8bb0

Observation 2f716131-ba4c-46d9-b828-528614592b22 · outbound

This paper cites Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Gen2Sim: Scaling up Robot Learning in Simulation with Generative Models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.797306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:ab93e7de2f41ed54286b8eac2436b8927285254384c55240b3e2599e5138b65e

Observation 78e2d6b0-013c-4fb4-b4a6-1452e4df1e5d · outbound

This paper cites DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DROID: A Large-Scale In-The-Wild Robot Manipulation Dataset

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.849645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:eb283b02dcf58225d6a14185cfb337a42496f44742b76373b1790797bb9472b8

Observation bfa1d90c-26cf-42e0-ade1-c92096b052c7 · outbound

This paper cites Learning to Act from Actionless Videos through Dense Correspondences.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Learning to Act from Actionless Videos through Dense Correspondences

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.801872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b73b0816d1367570a50fcdaa93f4a2ea31229b393aa96be346a8db4c710bcadd

Observation 4526ae40-b113-476f-a080-709b1c6f6d87 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.870240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:e01d841c1a30dd6ea8289cd1da7b264e5326f98bb42d6596899258ba41cf9efb

Observation 2d13c7cd-c211-408c-b5b6-329059d663e1 · outbound

This paper cites Masquerade: Learning from In-the-wild Human Videos using Data-Editing.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Masquerade: Learning from In-the-wild Human Videos using Data-Editing

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-29T02:04:55.330788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:7ec9e485fcc95a0462787d75d1dbd49e9b3eaacac940f3af94ef42073050668e

Observation 05ba843a-67e1-42fe-b6b5-8988a52f79d1 · outbound

This paper cites JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.853124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:802c981e651779e09836e7b74b9d2a07815b5e51ef034c344bb213ef8510e8ef

Observation 332cbd26-49ec-483f-95b7-4ef5908eeb2a · outbound

This paper cites Flow Matching for Generative Modeling.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Flow Matching for Generative Modeling

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.856143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:347ddadc0b82f8ae9fb5fb0311e162281938fdea0daa59834c1f1fc5b18c5664

Observation 66320006-eb3b-4b5a-9130-802212c47e73 · outbound

This paper cites Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Simplifying, Stabilizing and Scaling Continuous-Time Consistency Models

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T10:26:23.772771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:6ac8f66fb0f3ebfb5ef5d5fa9a7222c344d5a5507071289f6e0918c793d9b231

Observation 8361f674-f046-4f97-8fa5-558fd3b4f82d · outbound

This paper cites DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DPM-Solver++: Fast Solver for Guided Sampling of Diffusion Probabilistic Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:54:11.877839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:313c4901252dafa26f7d0ad97aedf4f44eb15d2ab2e7ec9b802472dcaa411d9d

Observation b8f0bf8d-4b90-4ec6-9ae1-bb888d1edd4d · outbound

This paper cites Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Beyond Flat Text: Dual Self-inherited Guidance for Visual Text Generation

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.902953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:0ffcd1e77e5b4f8f953d91f00c8e7b47cf3ba1e43ebdde7c0dc49f143c1e09aa

Observation 5e221f0a-481d-4cc8-9967-be2c3009115e · outbound

This paper cites Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Glyphdraw2: Automatic generation of complex glyph posters with diffusion models and large language models

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.711833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:aa47e85d2488cb9a5be0e63c96a7ba9e7c4ebb2bc246d1b304189d703f2b82e4

Observation 1ee0b7ad-b2dd-4f7b-95de-f11021a4e47e · outbound

This paper cites ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling ShapeSplat: A Large-scale Dataset of Gaussian Splats and Their Self-Supervised Pretraining

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.756011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:9a6e5b1949a2ece5a3a673759c28ebd6c003c9ed7fc06c32ed4e3dcc1c4839ca

Observation 7aaf1ffb-0a44-4032-9285-083dc079d353 · outbound

This paper cites arXiv preprint arXiv:2508.15772 (2025).

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling arXiv preprint arXiv:2508.15772 (2025)

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:21:27.647034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4b2ee99b6ba3eff7da1ea5c95574d874cc8e0f19b0c32d60b6e394adcfae2bf5

Observation 97341eec-7b27-4499-8664-e3b19b332603 · outbound

This paper cites Video generation models in robotics-applications, research challenges, future directions.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Video generation models in robotics-applications, research challenges, future directions

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.665259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:1e051c0a01dc945f0a343a571135a4cc608f735581939a0dbfef035e38e833ad

Observation 92504bca-38a7-4ee8-9596-3c11d3bd1cc6 · outbound

This paper cites PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling PhyBench: A Physical Commonsense Benchmark for Evaluating Text-to-Image Models

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:29.045357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:062ce1e40053d500d25478a2e52e05bc267e92670b78caa168f44f5878285da2

Observation 8a66066d-bc72-4a5c-b6a6-acc8a886cffe · outbound

This paper cites Spectral Normalization for Generative Adversarial Networks.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Spectral Normalization for Generative Adversarial Networks

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.726775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:9f9871a4ef9e920c12832d02019fda99a495fc69f3362597c3f85c1fd32cb9cd

Observation 886184f1-9194-4657-92a0-3c4e4e7f6c31 · outbound

This paper cites Chong Mou, Yanze Wu, Wenxu Wu, Zinan Guo, Pengze Zhang, Yufeng Cheng, Yiming Luo, Fei Ding, Shiwen Zhang, Xinghui Li, et al.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Chong Mou, Yanze Wu, Wenxu Wu, Zinan Guo, Pengze Zhang, Yufeng Cheng, Yiming Luo, Fei Ding, Shiwen Zhang, Xinghui Li, et al

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T09:59:02.389847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:6d0a1db756cecbe2ddc8a7517f101fc62cfad02eee90654b5daf87bd93d35832

Observation 5a358dda-d8c3-43da-ba5e-d634c1b37138 · outbound

This paper cites Transition Matching Distillation for Fast Video Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Transition Matching Distillation for Fast Video Generation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-13T01:17:54.723797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:20133796582013058a7f94aa8b6b3d74bd8fea48807adbb5471de4f6a583c016

Observation 509556b7-9c0b-4a52-8ea2-ae396f2bbae1 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4898c649aeb38ce617800012d03774fff13a17918a193a4ac5a695bccdc942a3

Observation 9629cf66-3253-4b46-be31-38f67d76a7f7 · outbound

This paper cites Improving robotic manipulation robustness via NICE scene surgery.arXiv preprint arXiv:2511.22777, 2025a.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Improving robotic manipulation robustness via NICE scene surgery.arXiv preprint arXiv:2511.22777, 2025a

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:29.026921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:0b2f1e1ee6dc22776961badad2013e4dbc473b0fbf44cfd42b280a57b9506852

Observation e69e67d4-846e-4860-a228-ef518983a16d · outbound

This paper cites Scalable Diffusion Models with Transformers.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Scalable Diffusion Models with Transformers

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:16:29.030285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:241d4a76f0a12421781cab2ca6a9036e5e52847e6a0cf15229f8f4b9f5563903

Observation c9b1d423-06c8-44bd-ba47-65329be4c6c4 · outbound

This paper cites Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Dustin Podell, Zion English, Kyle Lacey, Andreas Blattmann, Tim Dockhorn, Jonas Müller, Joe Penna, and Robin Rombach

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:29.038040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:d261e39e0daa770dec09c9363818332e8d68d222c503b5af29f3a219f55b456e

Observation 1f61d26d-0033-46d0-9a79-5bc2c92bcb54 · outbound

This paper cites Aligning Text-to-Image Diffusion Models with Reward Backpropagation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Aligning Text-to-Image Diffusion Models with Reward Backpropagation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:29.010694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:5be6f0680eba76a674fbcd28fd0493e22e00cf23145d4b46bdc365a91094ee2e

Observation 34aeac9d-c9d5-4433-9b25-297fb8bb6639 · outbound

This paper cites ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling ART: Anonymous Region Transformer for Variable Multi-Layer Transparent Image Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:29.019219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:8f6ab0a0d76522b7a719fb7cdf882d6fba83f3edee23ae4de5441121ddcdfc9d

Observation 226722f2-0983-4fef-883d-c2374909187d · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:29.041921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:571bcc257a34c82fba13c6b12dfb38363b0bfc12bf831491308bd1001a0604a0

Observation 94712609-16cf-45de-a0bf-fb3541bfaba2 · outbound

This paper cites Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Unsupervised Representation Learning with Deep Convolutional Generative Adversarial Networks

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:44:09.291693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:5345084c9c22249c3243d735ca193a9c504ff5b0c039bc2345a8cfa0f05f780b

Observation b9251e50-cd0d-4954-b515-fb953e1f23ae · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling High-Resolution Image Synthesis with Latent Diffusion Models

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.733469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:03db7a0efa9f9d47a2a377b90a7ae531248df91df9a9183d3bebaced7361ddfb

Observation b0aaa0d1-70e1-4ebb-935b-6617344255a9 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:21:27.719080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:41acef89c072d320fa0efba9e29168815d667c1111bb34e6e2da3e1a0aa2eb0d

Observation 941a017f-50aa-44e5-88cb-2c735dc369ba · outbound

This paper cites Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, et al.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Samin Mahdizadeh Sani, Max Ku, Nima Jamali, Matina Mahdizadeh Sani, Paria Khoshtab, Wei-Chieh Sun, Parnian Fazel, Zhi Rui Tam, Thomas Chong, Edisy Kin Wai Chan, et al

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.990967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:f0e457a78f9ccada6c09d3775f2a90083a50bae1a77a9b16d664ecb3afd9e0bf

Observation 3ed3727a-2f08-44e8-ab83-1edfeaa96ad4 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Proximal Policy Optimization Algorithms

Reference 68

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.998792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:efd55cd6b5ecafdc11075f640d4461dccfab70df5e350d5e1db7d7049d8f5b82

Observation f0180740-5d93-476b-8863-53d1152ef251 · outbound

This paper cites Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.982692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b7587b4c8692aafce463c581069f05d0de7780e05eb518382b3bdcd546d04c13

Observation 7fbb1b32-12e5-4863-a75f-5f5676e0699e · outbound

This paper cites Post-training quantization on diffusion models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Post-training quantization on diffusion models

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T09:59:02.393103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:f08949cb99006736e5a71be1274325e1e1245eaa891d97606992fe30394e62e8

Observation 7be5592e-5311-4bcc-9e7f-931f12e1a5ce · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.956821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:954521a9509950c4d723c3c5b62cb95dc14a7a4315b097f34cc7070ba02e2a99

Observation cc526aa0-a7ab-4acc-860c-85689abef157 · outbound

This paper cites Imagharmony: Controllable image editing with consistent object quantity and layout.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Imagharmony: Controllable image editing with consistent object quantity and layout

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.979081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:1a4864f318dc76d472e04cc995f81383a8a20e2976a1af50c762d62a2939c833

Observation ddabf6d7-b603-459a-9223-e82c2a12e401 · outbound

This paper cites Videovla: Video generators can be generalizable robot manipulators.arXiv preprint arXiv:2512.06963.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Videovla: Video generators can be generalizable robot manipulators.arXiv preprint arXiv:2512.06963

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.895006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4f7bdda8d89020a93d64b430d5ec0a3240ed711062fc0842e875b4e087b57314

Observation f9fe99c4-1a9c-4d40-9b73-f5f111a5ae11 · outbound

This paper cites Latent diffusion model without variational autoencoder.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Latent diffusion model without variational autoencoder

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:29.006820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:445d06ba60b818d3f0e826f386af26bda052ab2b66e66c751b5a5bb37ca7a386

Observation 4f2e2728-2413-4fcd-b7ec-ed7246988402 · outbound

This paper cites Chimera: Compositional image generation using part-based concepting, 2025.https://arxiv.org/abs/2510.18083.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Chimera: Compositional image generation using part-based concepting, 2025.https://arxiv.org/abs/2510.18083

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.964703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:59c70daf95c26df1475d3bba8dde112c86e56509bc6cfacacab0bf03e4497396

Observation 73f24d41-de09-4a7c-98b4-fd80e0734322 · outbound

This paper cites Mitty: Diffusion-based human-to-robot video generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Mitty: Diffusion-based human-to-robot video generation

Reference 76

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.971485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4ad73ad9845ce74e46973727f75435f9bd8fcf17eab18481cae6fd8da978fb21

Observation 2c44da9a-67f3-44f3-8d4b-dda5954a8faf · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-12T10:16:28.878385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:f3ff8325ae2b4024453f4d639e9f383ae7430ca69f322eec7192a5096030a3f7

Observation d0ee0a63-644a-47e0-b0fd-c03b2c37044f · outbound

This paper cites Understanding Generative AI Capabilities in Everyday Image Editing Tasks.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Understanding Generative AI Capabilities in Everyday Image Editing Tasks

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:27.680170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:88d70d3be8d17a13b05b7fa7dcf85e4f76135bd6091ce4995db9ec966e70db33

Observation 7c6420ec-3284-4d02-98d6-12d4e56b3f0a · outbound

This paper cites InstantCharacter: Personalize Any Characters with a Scalable Diffusion Transformer Framework.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling InstantCharacter: Personalize Any Characters with a Scalable Diffusion Transformer Framework

Reference 80

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.874266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:a4239dd63f33307c0971d414435bf0860fd8df8a08fecb4a28792267fd92c949

Observation 3e575da1-abaa-4afc-bf89-fbaeaae2d3b6 · outbound

This paper cites Diffusion Models Are Real-Time Game Engines.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Diffusion Models Are Real-Time Game Engines

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-16T12:04:44.468386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:4b420beb31689e4e6ad2530a92c22e04f63076eb04ba89f326af08d69d8ec7cf

Observation cca2e91f-52bd-46cc-976f-13d15fb18276 · outbound

This paper cites DataVisT5: A pre-trained language model for jointly understanding text and data visualization.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DataVisT5: A pre-trained language model for jointly understanding text and data visualization

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-27T09:59:02.399631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:d9ada916d9c85f1565528a0cf6de2bd64bf02e9dcc81a5db2b81a840a4940c1a

Observation e09ee0c7-9e2d-4bd0-99d3-5c2fb386962f · outbound

This paper cites Textatlas5m: A large-scale dataset for dense text image generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Textatlas5m: A large-scale dataset for dense text image generation

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.920683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:e2536cd9341565aadae76c08e3bd8f99d6478e050da07773e6af0bfff8d323aa

Observation 44fb26d3-5e15-4b05-a866-72752536c2b9 · outbound

This paper cites Growing visual generative capacity for pre-trained mllms.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Growing visual generative capacity for pre-trained mllms

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.929577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b63bab622130fbfdaa85d34c0925c52adae2478151d84cdf8a3df2b585dc046b

Observation 6b2b842f-6b4f-4a34-9a2b-3a89513fc44d · outbound

This paper cites UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling UniGlyph: Unified Segmentation-Conditioned Diffusion for Precise Visual Text Synthesis

Reference 85

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.886236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:87cd632831f1c1a735a1e7f132bb8f8fe9f4c71baa6d6712591cf2b97aa92a95

Observation fabfa5ae-5720-45c9-9351-a400265671c0 · outbound

This paper cites Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.907258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b83fdf0cf930f45472ef034f0e0aec2553460d332c114b099b62efc9e0461354

Observation dfff29d6-041f-402c-a281-dc240ec41730 · outbound

This paper cites SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.702876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:996924dde3d4eaa8257dbbd24f067d2d1ccdb5d07620b4719074117a9b4b6886

Observation 0b4b0d97-2d95-436b-9de7-812c69784375 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 88

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:21:27.637167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:003eeac270ce97ac059d4387e70b3bcc581701bbc18623d6dc88f29a7a1b445e

Observation d4a48a04-99d4-4dbf-ae14-f4c624529792 · outbound

This paper cites VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling VisionReward: Fine-Grained Multi-Dimensional Human Preference Learning for Image and Video Generation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:49:14.981353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b318ee99d08ed76e1e9dfec87d4e764fb9f3dc61c5375322348efc8e65c7a457

Observation 0b35d0d6-5578-45aa-9f45-8d3b6fb2350c · outbound

This paper cites DanceGRPO: Unleashing GRPO on Visual Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling DanceGRPO: Unleashing GRPO on Visual Generation

Reference 90

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.882129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:19508e7866594d852c2fbd18c406ca7d612f4b567cd2f34e6e40abe8f08c7529

Observation 9438478e-398b-4ed5-ba4d-ddb8b5d3cf59 · outbound

This paper cites Can understanding and generation truly benefit together–or just coexist? arXiv preprint arXiv:2509.09666.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Can understanding and generation truly benefit together–or just coexist? arXiv preprint arXiv:2509.09666

Reference 91

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:28.832206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:9b22c7a6b0d29eb634360414d27b8817408e88e4c034f034a975523800586e8f

Observation 3159d70f-e105-49d8-8a7d-c3f839fd79c7 · outbound

This paper cites EditWorld: Simulating World Dynamics for Instruction-Following Image Editing.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling EditWorld: Simulating World Dynamics for Instruction-Following Image Editing

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.792859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:bf4f0b69483fffa0fbd7825b1018e7a4baa852dc0984ee88b92ae8f95975e678

Observation c613c27c-4828-470a-8e0a-0d8b64c427b1 · outbound

This paper cites Rechar: Revitalising characters with structure preserved and user-specified aesthetic enhancements.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Rechar: Revitalising characters with structure preserved and user-specified aesthetic enhancements

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T04:55:12.101983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:70eb0dbb0f9a7c88492e889159199bff1f0c87ac4704c51e430cc80ca83ad7a7

Observation 0a1f991b-a8a2-4790-984f-a3cad0cb0bcf · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:17:45.693018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:320c38b18726274734e52f9e6e9930216a4775d08bfe0adf02301f4c5a789168

Observation fbc62f89-28e5-4d7d-816c-a386edc562cd · outbound

This paper cites ReasonEdit: Towards reasoning-enhanced image editing models.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling ReasonEdit: Towards reasoning-enhanced image editing models

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.842566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:3872ae246fceddbeb0b69f401df012a05efd801c7790853782967dff5b3e5b42

Observation 5a8b9bcf-f36c-4e10-8bdd-fa3f024ee12e · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 96

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T15:09:37.504223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:c0fd3373380855792ef403af4ded4075d666bc6bbef269121e91deeec0e3ac2e

Observation 90cfe21e-7a94-47fb-bc7b-4eb3d23123f4 · outbound

This paper cites Mira: Multimodal iterative reasoning agent for image editing.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Mira: Multimodal iterative reasoning agent for image editing

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.825816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:b1d85327426fd520c82a2266a1480450136dee95fb5175757f30250fb55f3976

Observation 90e92c45-3511-40f5-a257-f526b0da77af · outbound

This paper cites CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.822243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:f4a71f505390403db6e87dc7d046f653861fe8364f6f2af0e627a165eff7a307

Observation 6968aefd-ff77-4704-bc2f-b7ed6a2fedfd · outbound

This paper cites Lvmin Zhang, Anyi Rao, and Maneesh Agrawala.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Lvmin Zhang, Anyi Rao, and Maneesh Agrawala

Reference 99

Resolution
verified exact
doi, observed 2026-05-09T04:55:12.096691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:837c4fbb7fcdd541a8af5affab90a785d0070a63b8f30886aec4c7e3c3141378

Observation e7ac35e8-e49c-4254-bd97-acd4e92c6ebd · outbound

This paper cites Make geometry matter for spatial reasoning.arXiv preprint arXiv:2603.26639.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Make geometry matter for spatial reasoning.arXiv preprint arXiv:2603.26639

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:16:28.814775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:2166cb879326e3750f05057c8e87af0bb6b20b992abe0fbcdacfa446111b4b84

Observation 4d3a7dcd-1803-449a-a1af-af0f7e6fb28a · outbound

This paper cites Text2Layer: Layered Image Generation using Latent Diffusion Model.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Text2Layer: Layered Image Generation using Latent Diffusion Model

Reference 101

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T10:16:28.829061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:7ff8eb1386bac689a7939bc54e55ee2db39e1a35e4b4f36fc7211e18962abb44

Observation c6aaefde-9015-4dee-a576-e9a405884fbd · outbound

This paper cites Diffusion Transformers with Representation Autoencoders.

Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Diffusion Transformers with Representation Autoencoders

Reference 102

Resolution
verified exact
local_arxiv, observed 2026-05-12T10:16:28.846706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-07T06:38:04.459129Z digest=sha256:ba977e7b88d51f424793bb8030c7befec5fa672fff285f85c9f603fddf577e8b

Pith citing papers

Observation 5fcb7a34-a8e5-4beb-bca8-f03a78139923 · inbound

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors cites this paper.

WorldReasonBench: Human-Aligned Stress Testing of Video Generators as Future World-State Predictors Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:11:25.214138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:39:01.254916Z digest=sha256:3afe71e5c514385bbbf6501898d0afe22745945b6514854e50e1926da398e24b

Observation 0c507a02-8afd-4c45-86f5-7f4b4afe0d8d · inbound

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping cites this paper.

Power Reinforcement Post-Training of Text-to-Image Models with Super-Linear Advantage Shaping Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-12T07:16:29.508326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-12T03:33:40.994346Z digest=sha256:f80ed99f6244a7655a84403531df18cd9c2871939ca00e827a9f7e9ffc8db73b

Observation 84a93aff-34d0-4f45-9b15-b5fc59da6f07 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.592635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:cb8ffb2fe611a81570fad060f97511cbce013d62395ea5dda62852c5085060e6