Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T08:15:58.020894Z
Paper Citation Record · LEDGER
As of 5 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 6 inbound Pith citation observations for arXiv:2605.04128.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T08:15:58.020894Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-02T06:14:03.146944Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-03T15:58:37.266385Z
100 of 105 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3abe9105-13f4-480e-8dd4-3f977d36dbce · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a22bd4ad-626f-4545-bf4e-6cd6cc98f565 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Recammaster: Camera-controlled generative rendering from a single video
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 64c01157-ef8a-4e39-a41e-63fbf52398f0 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen3-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ae098007-f369-4453-9a7d-ae96f6de0f4b · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3fd5c8a-be96-4a69-b054-931b93173e49 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a089d5d9-46e3-4be1-a4fd-bb4d3709d8d5 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Black Forest Labs
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03720101-31da-4fc5-80ff-83fbdda66462 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Blender
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 270bbc30-87ec-42e5-b22c-e95d83ed4fcb · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Instructpix2pix: Learning to follow image editing instructions
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bb54c90-c8f8-4707-b918-9f234e0b2033 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Genie: Generative interactive environments
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7880d8de-e98a-4bba-aade-79867b5a9ee8 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d14a179a-f6a1-4277-acc1-ad0bbe483241 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4ba35277-8856-4916-9c57-f19ebcaa5e0f · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SAM 3: Segment Anything with Concepts
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ffc603a-38dd-46c1-866a-ac08214f0e5d · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation WorldVLA: Towards Autoregressive Action World Model
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4e8432ae-cc1d-43e9-9664-2f81550449f2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Matterport3D: Learning from RGB-D Data in Indoor Environments
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3f9af78d-bcf5-484f-980b-82f84f2f67e2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OneIG-bench: Omni-dimensional nuanced evaluation for image generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4a0ba0d9-10f7-4c95-acd1-7ed79dbb5619 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Textdiffuser-2: Unleashing the power of language models for text rendering
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b16c8be4-5b30-4d38-8b19-c57d88bc2e15 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 007c3f1e-b5e9-4c91-8461-e47baac70a0b · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7d3d69fc-9252-4273-828a-e0e44703fc48 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 18a1bc22-24a9-4814-99ea-7905e4de24da · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 37ffc3e4-3e5d-4187-b06b-3593aa08e398 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 342cf724-f227-4e54-976d-80d1d7540789 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b186d508-3ddc-4e75-ae18-7e893b9d9193 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8317616b-b481-4b10-b5f2-0e15bf1109d2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 389f1f92-6c60-4f67-a9ce-bb8d79c483e7 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0e7da6d1-7b74-4ab6-8b69-c80d7e5ebb7f · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scannet: Richly-annotated 3d reconstructions of indoor scenes
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b0b29226-4580-4515-ae56-f2842bc9ad10 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 806fb2ae-0592-40f0-b2e5-2f616bbf0f98 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec6ae8e5-2f42-4765-9cb8-ea6ce7d5f563 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10543328-5c79-46aa-a1b0-ec8388a19fb2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Blink: Multimodal large language models can see but not perceive
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c6aac27d-9d4c-4929-92e3-db0e508602b1 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seedream 3.0 Technical Report
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4111d9bc-256b-4d43-a35c-91b645b73ff9 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4b4a24b9-0b12-4f9a-9a48-67e7e706e0e4 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation A new era of intelligence with gemini 3
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fe7e7989-68f3-4bd4-bee8-472161cfe817 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Nano banana pro
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb6c4115-10e5-438b-bfec-1dfca8ebbff4 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos.https://deepmind.google/models/veo/
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8855698f-59fd-49e9-b076-6c8e4855f8ee · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0bbf9d77-95ae-4471-bf5a-682e72a57b54 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 589cc1fa-1ca8-45d9-9b98-1114c8dd1e88 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Musiq: Multi-scale image quality transformer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation caa0a05a-182e-4f13-a584-e2bb9876a83f · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OpenVLA: An Open-Source Vision-Language-Action Model
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ccff3446-873e-42e5-a248-841acf9b6afe · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 54188d15-69ed-411c-9afb-02649c6ec407 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7b3e1358-e2fc-4ec1-b2f4-460e597e2866 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation FLUX.2: State-of-the-Art Visual Intelligence.https://bfl.ai/blog/flux-2
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c58781c0-f101-4a58-83d1-a3f531ad8028 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Easier painting than thinking: Can text-to-image models set the stage, but not direct the play?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5a42f053-6991-4b1e-bd10-9cf75fee6109 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d938c7bb-05eb-4152-a75a-b14da56f16f7 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d5822dcc-0b2e-4193-980d-3f54d4b42508 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flow Matching for Generative Modeling
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ff7d5e66-e80b-40d0-85d3-0683a070c14e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flow-GRPO: Training Flow Matching Models via Online RL
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9fd3a95b-8c5f-49cf-ada2-17da193727e2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e2da9cb-cd45-420e-82f3-c0825f1a5675 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Step1X-Edit: A Practical Framework for General Image Editing
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a9284257-c9eb-44a5-b624-509ae8ca760e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? InECCV
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2377bf2d-07be-4cda-960a-5db40982e10e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation adb1cb8b-c251-4eb3-a902-8b23a923eb0b · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9b322a8-debc-4ba4-aa14-534b8d053d6e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Editscore: Unlocking online rl for image editing via high-fidelity reward modeling
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f1465833-6c47-498f-ad39-b29fc13a781c · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation X2i: Seamless integration of multimodal understanding into diffusion transformer via attention distillation
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0398a180-9b34-48cd-ac20-d9fdbe13f509 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation 3dsrbench: A comprehensive 3d spatial reasoning benchmark
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ad32dff7-2ede-404e-825d-ea9f765d3b4b · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Hpsv3: Towards wide-spectrum human preference score
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 61cc88c1-9d64-43fc-b730-0425291c125c · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation completely blind
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ee9ff11b-337c-405b-a205-7711e34501da · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Chatgpt.https://openai.com/blog/chatgpt/
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cee8d0e1-3709-4160-9311-97da1a5458b7 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT Image 1
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb99cbff-27e0-4d29-857c-e6efa5a5af7a · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SpaceR: Reinforcing MLLMs in Video Spatial Reasoning
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 83f6a747-2e6c-4e31-aa26-ee2fc864882e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pico-banana-400k: A large-scale dataset for text-guided image editing
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 46aa6c90-740d-4132-a9ad-1e797b3aa0d2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Camedit: Continuous camera parameter control for photorealistic image editing
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fc0d46ba-b7d8-4534-b26a-acc4cf649144 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vincie: Unlocking in-context image editing from video
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5cc168a-26c0-4089-b565-548536802468 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Ultrapixel: Advancing ultra high-resolution image synthesis to new peaks.Advancesin Neural Information Processing Systems, 37:111131–111171
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f9159f5d-dccf-4a8f-b52d-08e6b4d8bce4 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f66d4e37-a400-4982-a2fd-b556c46054ab · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 80f1f1e9-185d-4525-bd79-a311c324eebd · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 77cabe27-51d5-40c4-aa38-b3a2361805c2 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b016ee3f-62b0-41f3-8212-62db4d9b106a · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 331279a3-f665-44ca-ae9d-d49469bd8e95 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8fb70b1c-f0f8-4f0f-9312-a5fdec7c5155 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation LongCat-Image Technical Report
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 436d8de4-99b8-4d4f-a188-1adc6043c2ae · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Advancing Open-source World Models
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c5ca38db-48af-4e58-87e0-116bba08dba1 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Firered-image-edit-1.0 techinical report
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a047d48-282c-4393-929b-712f74b3d7de · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 544d6ec0-1624-4948-9a40-0194af737028 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cc19099f-af49-4ab4-aafe-65c0ab765637 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vidu: Ai video generator
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1939d2ab-5fdc-419b-a1e5-f76610f2ee9f · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Wan: Open and Advanced Large-Scale Video Generative Models
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51c6301b-3c18-4224-b18c-e901feadcb94 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Exploring clip for assessing the look and feel of images
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 62c2ceba-f59d-49ab-8d96-4d3dc75b34fc · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vggt: Visual geometry grounded transformer
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 03bec716-023f-4eae-bdb0-2b5a876067e0 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e3eddba3-60ac-422b-a8cf-c18f06e73ced · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8f05874-1ceb-400b-8df4-e660e8ca29c1 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen-Image Technical Report
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ca22cb71-1ee9-484d-ae73-97c5cd52e08f · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71eb7b9a-4e8c-4d42-bb1e-43099e67967a · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Alvarez, Jun Gao, Sanja Fidler, Zian Wang, and Huan Ling
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation de0a1c2b-5f2e-46ad-ab30-fb14621b935e · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Grok-1.5 vision preview
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9605d07a-404d-4606-8547-0ce231fe3026 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Dreamomni2: Multimodal instruction-based editing and generation
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation eaf52597-4aa9-44d8-b578-cada8fdf0b21 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Omnigen: Unified image generation
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47fafb5d-8773-4a65-a555-1a1003b673a0 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation fb2ba9fb-ca4c-4815-b554-783942a8e864 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen3 Technical Report
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f7333f8b-7233-4cd0-82f0-ddfb978dc482 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Thinking in space: How multimodal large language models see, remember, and recall spaces
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3ce33607-ebf0-4035-a00d-70998431b80c · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Thinking in space: How multimodal large language models see, remember, and recall spaces
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f72dfe91-3bbf-45d3-ba9c-7bd8c2e6dc9d · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cbfcce05-18cd-497b-85f1-4b8ddc6a3cff · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ImgEdit: A Unified Image Editing Dataset and Benchmark
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 38df0576-8768-4231-942f-888ddd89cf95 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9f6c08b5-40e4-470d-8125-f9384766fe12 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scannet++: A high-fidelity dataset of 3d indoor scenes
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1df1c87b-61e2-4aee-88fa-de82558a3546 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Anyedit: Mastering unified high-quality image editing for any idea
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 978e6a56-9ef4-4171-8265-f8268c3936df · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Magicbrush: A manually annotated dataset for instruction-guided image editing.Advancesin Neural Information Processing Systems, 36:31428–31449
Reference 101
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ab95a6e8-f54d-41a3-afe0-f89faf92c57a · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Bamboo: Building mega-scale vision dataset continually with human–machine synergy
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3e43ecc3-b147-4a01-9e6f-b3443548d234 · outbound
JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer
Reference 103
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 305321f0-c390-4fac-b274-50ddbeb32e94 · inbound
Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 91e3a084-0cfb-4708-aa47-020927d0415e · inbound
Qwen-Image-Flash: Beyond Objective Design JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 372dea67-29d7-4177-82ad-c6ac9e1b8d89 · inbound
DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dd050f85-7c63-4513-bdc2-6e6af8183190 · inbound
Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 118
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d96657bd-7282-46d6-9cc4-98e0a09575bc · inbound
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 98
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1a5d00b-283e-4073-964a-8a37a63f970c · inbound
Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.