Pith. sign in

Paper Citation Record · LEDGER

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

As of 5 August 2026, this Paper Citation Record lists 100 of 105 outbound references and 6 inbound Pith citation observations for arXiv:2605.04128.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.04128 v2

Coverage vector

measured 100 of 105 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T08:15:58.020894Z

measured 106 of 106 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:14:03.146944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-03T15:58:37.266385Z

Reference resolution

100 of 105 outbound references displayed

  • verified exact55
  • verified fuzzy43
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3abe9105-13f4-480e-8dd4-3f977d36dbce · outbound

This paper cites GPT-4 Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.733582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:cdd45057e031e3f5a50619d8717f4df7d47e066cb61b7f46925ce24099289392

Observation a22bd4ad-626f-4545-bf4e-6cd6cc98f565 · outbound

This paper cites Recammaster: Camera-controlled generative rendering from a single video.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Recammaster: Camera-controlled generative rendering from a single video

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.894886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:536e4e8a29025860d1a9bbfe3ed0db6e636f7b35be84de4e2affb74c8bb2ebaa

Observation 64c01157-ef8a-4e39-a41e-63fbf52398f0 · outbound

This paper cites Qwen3-VL Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen3-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.677383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:a9a86a774a749c0f870421aa3ecbc39fc2b879c099b9556d9124e25b70349a1b

Observation ae098007-f369-4453-9a7d-ae96f6de0f4b · outbound

This paper cites Qwen2.5-VL Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.720463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:0577b889eb160f3355356d23d00f8d74756f26841fa41b03be8b36943dde0815

Observation e3fd5c8a-be96-4a69-b054-931b93173e49 · outbound

This paper cites ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ARKitScenes: A Diverse Real-World Dataset For 3D Indoor Scene Understanding Using Mobile RGB-D Data

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.723702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:12659c3d567a13775576cc1d31aa27bdb74fd7ba75cd040161817ae135854db5

Observation a089d5d9-46e3-4be1-a4fd-bb4d3709d8d5 · outbound

This paper cites Black Forest Labs.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Black Forest Labs

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.892923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ebc4fa6c8b9be46314af869c78ab15d1b469918bba9d54211b05896ea490102c

Observation 03720101-31da-4fc5-80ff-83fbdda66462 · outbound

This paper cites Blender.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Blender

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.882455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:dd25d264052b2b6f45503f3e0165f47cec9de740fc1e911b7b8db3888d3b6509

Observation 270bbc30-87ec-42e5-b22c-e95d83ed4fcb · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Instructpix2pix: Learning to follow image editing instructions

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.887001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:06469c3fcfd76de0fe0b1ed0046e0a9c0eb1969854d24b94be87e7650c574d85

Observation 9bb54c90-c8f8-4707-b918-9f234e0b2033 · outbound

This paper cites Genie: Generative interactive environments.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Genie: Generative interactive environments

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.900167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:2b557916440232e157dbbb94dea9501d9f5fa5a560a8e8cc770e03e155a644b1

Observation 7880d8de-e98a-4bba-aade-79867b5a9ee8 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.709965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:dfe2a201abca7052e66f1945aea12ea03dcc8b7a09f8c0cac33cdca1c962085b

Observation d14a179a-f6a1-4277-acc1-ad0bbe483241 · outbound

This paper cites ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ArtiMuse: Fine-Grained Image Aesthetics Assessment with Joint Scoring and Expert-Level Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.717102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:a6f28222e8606de02241aacc352c61adce3df7649fe1b1d933ca44f0fc2d01da

Observation 4ba35277-8856-4916-9c57-f19ebcaa5e0f · outbound

This paper cites SAM 3: Segment Anything with Concepts.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SAM 3: Segment Anything with Concepts

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.689803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:aa7d292a474a60fef3fe58b43763a43204113f0d87d5ae6fd509fbeaff8005c7

Observation 3ffc603a-38dd-46c1-866a-ac08214f0e5d · outbound

This paper cites WorldVLA: Towards Autoregressive Action World Model.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation WorldVLA: Towards Autoregressive Action World Model

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.611987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:51ba3b19ff122f0c7394949c48aac884d6cd85f3208115bcb0218e97dd572810

Observation 4e8432ae-cc1d-43e9-9664-2f81550449f2 · outbound

This paper cites Matterport3D: Learning from RGB-D Data in Indoor Environments.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Matterport3D: Learning from RGB-D Data in Indoor Environments

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.713521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:33eaad3b31583fe1238a1b2fa3f8175b373bc53b0570be39cefb1123f37eeb32

Observation 3f9af78d-bcf5-484f-980b-82f84f2f67e2 · outbound

This paper cites OneIG-bench: Omni-dimensional nuanced evaluation for image generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OneIG-bench: Omni-dimensional nuanced evaluation for image generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.811167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f2c1a100e6c40b9c1eef13ef349c9bd9439275f7cf5ce0ca909be2b5f4493e07

Observation 4a0ba0d9-10f7-4c95-acd1-7ed79dbb5619 · outbound

This paper cites Textdiffuser-2: Unleashing the power of language models for text rendering.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Textdiffuser-2: Unleashing the power of language models for text rendering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.880409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ffdac7fc8c72108bcd60c352ca6b74fca6c1a662f0c7eda43f18fae1993438a4

Observation b16c8be4-5b30-4d38-8b19-c57d88bc2e15 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.605080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c071cc485348109500fcd36ba6ea24d43af54e8ffef46667e249e18034167925

Observation 007c3f1e-b5e9-4c91-8461-e47baac70a0b · outbound

This paper cites Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pixart-σ: Weak-to-strong training of diffusion transformer for 4k text-to-image generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.884752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:a57d5ccf9b2f2db9a7b8187688bc4c23c48567dc4e6ef0f6c8ac00c02664aadb

Observation 7d3d69fc-9252-4273-828a-e0e44703fc48 · outbound

This paper cites Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pixart-α: Fast training of diffusion transformer for photorealistic text-to-image synthesis

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.898173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:022a78388d2061299ee58c9a100fcba379ac4cd6def34836857124885587ba4b

Observation 18a1bc22-24a9-4814-99ea-7905e4de24da · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.706959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:aa86900cb757c2b54d27c6ace3feda1fba0266e6521224465e8e776c6a75bf1d

Observation 37ffc3e4-3e5d-4187-b06b-3593aa08e398 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.622272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:afbc6a64b33928e6db040884980ba165916404aefc43476ac643311fc6fb4d9b

Observation 342cf724-f227-4e54-976d-80d1d7540789 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.736762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:7713a7ffcc98509087dd4cf9f2d168171c9e2b01d1cfc294f1fee72edf447dc1

Observation b186d508-3ddc-4e75-ae18-7e893b9d9193 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.668025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:43c0f7db15c8fc4bf0e1c9057f8eb6998b60197eddf5a29f63ae760ef9f289f8

Observation 8317616b-b481-4b10-b5f2-0e15bf1109d2 · outbound

This paper cites PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation PaddleOCR-VL-1.5: Towards a Multi-Task 0.9B VLM for Robust In-the-Wild Document Parsing

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.654846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:0f5e6f41a2e4af824a05dd38723ab2b06f5d6fb8885e019310674e6cd3673f6c

Observation 389f1f92-6c60-4f67-a9ce-bb8d79c483e7 · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emu3.5: Native Multimodal Models are World Learners

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.730252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:148c3b8834e36a21bdfc571f1a04c9324adadfa1d95ed0cf6d22851bfc2eb961

Observation 0e7da6d1-7b74-4ab6-8b69-c80d7e5ebb7f · outbound

This paper cites Scannet: Richly-annotated 3d reconstructions of indoor scenes.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scannet: Richly-annotated 3d reconstructions of indoor scenes

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.904627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:2cc44034c39b2c5ff86ad24f3896db85761bd3624479e51311ebc59c8e27d6cf

Observation b0b29226-4580-4515-ae56-f2842bc9ad10 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.664709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:1970354812d12ea5bdd3d3a7b8cc10fa30377975666925d1c35d0fcbdbb29534

Observation 806fb2ae-0592-40f0-b2e5-2f616bbf0f98 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.876093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:96145b7d89e79a71a4d542acea5f02fcb3e81adcdbdc66292b62a8c22452539f

Observation ec6ae8e5-2f42-4765-9cb8-ea6ce7d5f563 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.871983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ffee306920f246ae79b36bc62c5ef95e1ae8ca9a0b36a87be00ff35e549a2515

Observation 10543328-5c79-46aa-a1b0-ec8388a19fb2 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Blink: Multimodal large language models can see but not perceive

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.867843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ca24ad286828763c9dd61322e1bb21f916d5be0b00b88685aa88b10de33353b6

Observation c6aac27d-9d4c-4929-92e3-db0e508602b1 · outbound

This paper cites Seedream 3.0 Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seedream 3.0 Technical Report

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.670811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:cafece49ba95bb3267b4b83d6d092418f9c87c6c7d9b8de4ea4c03cb94c07078

Observation 4111d9bc-256b-4d43-a35c-91b645b73ff9 · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.658199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:ec61b16011139af4d874e86d6d331f282a427413402823a99c37bae226a9e5b2

Observation 4b4a24b9-0b12-4f9a-9a48-67e7e706e0e4 · outbound

This paper cites A new era of intelligence with gemini 3.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation A new era of intelligence with gemini 3

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.813154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:36a76a53bfd26ad68ce142bea292d4c5f190f7da14c3819d90ddb84e4b8d5e49

Observation fe7e7989-68f3-4bd4-bee8-472161cfe817 · outbound

This paper cites Nano banana pro.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Nano banana pro

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.859122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c04ad9e546bf15b8a3d705e546325250354bff787a29b7088cfbc97f593c510f

Observation cb6c4115-10e5-438b-bfec-1dfca8ebbff4 · outbound

This paper cites Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos.https://deepmind.google/models/veo/.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Introducing veo 3, our video generation model with expanded creative controls – including native audio and extended videos.https://deepmind.google/models/veo/

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.861234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:a2dd1470bccf2a29b12035d73a5bc83e69ec0e6cfa2717b52bf11bdbef22933e

Observation 8855698f-59fd-49e9-b076-6c8e4855f8ee · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.641695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:5ba0af1244a793d8ada86e0a98ca91ed490f4b1b9fe42ee285df24a4a63cdc5b

Observation 0bbf9d77-95ae-4471-bf5a-682e72a57b54 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.683500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:d8b5ba6be4a3ef835e57510b6fcc03906884b978d4e9c6a18ba8bd19354fad2f

Observation 589cc1fa-1ca8-45d9-9b98-1114c8dd1e88 · outbound

This paper cites Musiq: Multi-scale image quality transformer.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Musiq: Multi-scale image quality transformer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.857249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:61d504a342204516b479a7a758575fb65f7bbe60374ca5a671d0ca7325cce99f

Observation caa0a05a-182e-4f13-a584-e2bb9876a83f · outbound

This paper cites OpenVLA: An Open-Source Vision-Language-Action Model.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OpenVLA: An Open-Source Vision-Language-Action Model

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.637682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:8def6d4c662848ecfc6e00d7367943ed1b29b2c8781e134cd4fd307160692b22

Observation ccff3446-873e-42e5-a248-841acf9b6afe · outbound

This paper cites an unresolved cited work.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-21T08:24:52.808924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:4735728f9c13d8211100c42084b61e6cd981eb68623e29fe9bbc1de3b540a924

Observation 54188d15-69ed-411c-9afb-02649c6ec407 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.852835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:1f86d932689657faa749e4a0bc3f1985aa84db7f18d320e14a62c0a6f5166c38

Observation 7b3e1358-e2fc-4ec1-b2f4-460e597e2866 · outbound

This paper cites FLUX.2: State-of-the-Art Visual Intelligence.https://bfl.ai/blog/flux-2.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation FLUX.2: State-of-the-Art Visual Intelligence.https://bfl.ai/blog/flux-2

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.855011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:480ada38e7de08fbb3f732c59f0cf3c23e19e23f9e5286e6042dfd1e9c498c3c

Observation c58781c0-f101-4a58-83d1-a3f531ad8028 · outbound

This paper cites Easier painting than thinking: Can text-to-image models set the stage, but not direct the play?.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Easier painting than thinking: Can text-to-image models set the stage, but not direct the play?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.608614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:38a0ecc821f5ecc5a48487c23ec78f82b6d3b01d7f79a433a69c9ed57af246da

Observation 5a42f053-6991-4b1e-bd10-9cf75fee6109 · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:01:19.940440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:68071c9c447eb592347dd8e66c7d4e9b93b78316887e6fb6592540307d2330fd

Observation d938c7bb-05eb-4152-a75a-b14da56f16f7 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.758187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:b0d08d7a3372dfa65795937f4978f4c1e589015a0f3c6d33c833cf2f9ae2d485

Observation d5822dcc-0b2e-4193-980d-3f54d4b42508 · outbound

This paper cites Flow Matching for Generative Modeling.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flow Matching for Generative Modeling

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.618952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:bf7b3de0e33fa03d3e366e3bf55c895c7c521fb63d371fdc9ef512a0bfd793f4

Observation ff7d5e66-e80b-40d0-85d3-0683a070c14e · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Flow-GRPO: Training Flow Matching Models via Online RL

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.651577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:93a49c4f36a34622ca890370a20647d2ec1eafcacbb13ed2dcd37d81e7186289

Observation 9fd3a95b-8c5f-49cf-ada2-17da193727e2 · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.648398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:8610db71cd77852221a7ab6af03c905e15949a2e01f16619124c09737d1508a6

Observation 3e2da9cb-cd45-420e-82f3-c0825f1a5675 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Step1X-Edit: A Practical Framework for General Image Editing

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.789966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:acd30328854b7fd4e17c46d3dfe8d2e535c202e233a3018417437c6ca3bb8e88

Observation a9284257-c9eb-44a5-b624-509ae8ca760e · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Mmbench: Is your multi-modal model an all-around player? InECCV

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.863547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f53626d36366a9712cd8f40b87cdf6fb0404e4b05e383a68258752ec24d604b4

Observation 2377bf2d-07be-4cda-960a-5db40982e10e · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Ocrbench: on the hidden mystery of ocr in large multimodal models.Science China Information Sciences, 67(12):220102

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.865730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:bf8bf17ac4094b9b63d5dc6c57c40f0d962400d70b73d24d916ad9e718e3b08e

Observation adb1cb8b-c251-4eb3-a902-8b23a923eb0b · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.786620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:80feb38abc2e8b0cb6afaedce122d15ef82144df1bbc7e59ce42a0ec94829154

Observation b9b322a8-debc-4ba4-aa14-534b8d053d6e · outbound

This paper cites Editscore: Unlocking online rl for image editing via high-fidelity reward modeling.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Editscore: Unlocking online rl for image editing via high-fidelity reward modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.793973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:07e95cabc0480de1cac0a36ea327628b22a99eea8581db9906ea1107b18195f7

Observation f1465833-6c47-498f-ad39-b29fc13a781c · outbound

This paper cites X2i: Seamless integration of multimodal understanding into diffusion transformer via attention distillation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation X2i: Seamless integration of multimodal understanding into diffusion transformer via attention distillation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.869787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:077717fc21ed4c0ad3c6230997ed0b41a0281aaba2714f573ea8740c50c8f489

Observation 0398a180-9b34-48cd-ac20-d9fdbe13f509 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation 3dsrbench: A comprehensive 3d spatial reasoning benchmark

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.777451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:3280ff0912b00df3b8a88bdb845f2c58ca161b200b76f78d29d6792c507bc729

Observation ad32dff7-2ede-404e-825d-ea9f765d3b4b · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Hpsv3: Towards wide-spectrum human preference score

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.874076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:fbb367a2286f632e8b35ece68e70423b9b7d0a3ae0ceac5c1ed38b9e26a8c2fc

Observation 61cc88c1-9d64-43fc-b730-0425291c125c · outbound

This paper cites completely blind.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation completely blind

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.902378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:3bcd7b1e0b9f1660a41a13023f651d81c07f4f7b5f6153c439f3e0c1192c3cf2

Observation ee9ff11b-337c-405b-a205-7711e34501da · outbound

This paper cites Chatgpt.https://openai.com/blog/chatgpt/.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Chatgpt.https://openai.com/blog/chatgpt/

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.843975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:9e3a9661acefd06b60e767692db8669224c32d484163489d6fe1a69525db9b54

Observation cee8d0e1-3709-4160-9311-97da1a5458b7 · outbound

This paper cites GPT Image 1.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT Image 1

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.816986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:70637498fb69f4af7c196bb5bc892c396bf1a71c9f1d1e5a6ba686f283abe4af

Observation cb99cbff-27e0-4d29-857c-e6efa5a5af7a · outbound

This paper cites SpaceR: Reinforcing MLLMs in Video Spatial Reasoning.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SpaceR: Reinforcing MLLMs in Video Spatial Reasoning

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.771557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:199c8ccf5fe7262b10f89c029b4535bdb9fa05442b64372b366dab2dc589533e

Observation 83f6a747-2e6c-4e31-aa26-ee2fc864882e · outbound

This paper cites Pico-banana-400k: A large-scale dataset for text-guided image editing.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Pico-banana-400k: A large-scale dataset for text-guided image editing

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.768263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:94fc6c803944d3db5bf5025f6309404a895e10840a7ca47aa3179474deeebd3e

Observation 46aa6c90-740d-4132-a9ad-1e797b3aa0d2 · outbound

This paper cites Camedit: Continuous camera parameter control for photorealistic image editing.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Camedit: Continuous camera parameter control for photorealistic image editing

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.842070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:8232dc251c3f380857bfaab8f9995bdeb2d3af6fb7a53ac92a072f63a1a42336

Observation fc0d46ba-b7d8-4534-b26a-acc4cf649144 · outbound

This paper cites Vincie: Unlocking in-context image editing from video.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vincie: Unlocking in-context image editing from video

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.815140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:31a9bbd10b51ea94353445a54b88cafb03b4c040fcc3b17b8da03ecd72b0d914

Observation b5cc168a-26c0-4089-b565-548536802468 · outbound

This paper cites Ultrapixel: Advancing ultra high-resolution image synthesis to new peaks.Advancesin Neural Information Processing Systems, 37:111131–111171.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Ultrapixel: Advancing ultra high-resolution image synthesis to new peaks.Advancesin Neural Information Processing Systems, 37:111131–111171

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.846108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:d7319f8a48a24181b2b1577377154865dfb1092e996d067acd417fe76a55ffe6

Observation f9159f5d-dccf-4a8f-b52d-08e6b4d8bce4 · outbound

This paper cites Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Hypersim: A photorealistic synthetic dataset for holistic indoor scene understanding

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.850607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:aa3519f3303dd26d3699a1d264106bce39f1ab3f11901cb411719d5cb0c4e181

Observation f66d4e37-a400-4982-a2fd-b556c46054ab · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.835775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c6c8897aa7be1c5df53c9b96216af0d808327fb8c7e9ea653eaf981413db21c5

Observation 80f1f1e9-185d-4525-bd79-a311c324eebd · outbound

This paper cites Seedream 4.0: Toward Next-generation Multimodal Image Generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seedream 4.0: Toward Next-generation Multimodal Image Generation

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.739879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:3cc733c0fe53df808f41c19eaead72a9df82abd7aef334aabe7de8c6ce97b462

Observation 77cabe27-51d5-40c4-aa38-b3a2361805c2 · outbound

This paper cites MiMo-VL Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MiMo-VL Technical Report

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.700497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:0d345cadf90bc53b446f0b957eee8c6b7f18fe064b0363407f16ff20c44bafd2

Observation b016ee3f-62b0-41f3-8212-62db4d9b106a · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Gemini: A Family of Highly Capable Multimodal Models

Reference 71

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.680536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:fd6d31de09b8ec9a3ad6bec03619ae5c84bfa1793ca09b59b8c0fee0c64efd9a

Observation 331279a3-f665-44ca-ae9d-d49469bd8e95 · outbound

This paper cites Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Kolors: Effective training of diffusion model for photorealistic text-to-image synthesis

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.823078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:312cc2740fc2f3e4961a6342a007c42e58aa28b4e268f4611f24588ed4020025

Observation 8fb70b1c-f0f8-4f0f-9312-a5fdec7c5155 · outbound

This paper cites LongCat-Image Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation LongCat-Image Technical Report

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.745748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:48fccc11b0fd044d514e1f8d56dfc4f3266dbd8c32c54ed65ee827871943a908

Observation 436d8de4-99b8-4d4f-a188-1adc6043c2ae · outbound

This paper cites Advancing Open-source World Models.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Advancing Open-source World Models

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.742673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:30e914a1db0d32cc904fc2abeed884d85e56e44d44f6db23e019d947e8419f64

Observation c5ca38db-48af-4e58-87e0-116bba08dba1 · outbound

This paper cites Firered-image-edit-1.0 techinical report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Firered-image-edit-1.0 techinical report

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.761459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:406267d2c0758460bc2ac9ca8444398b05ec2cf941e4a7834c2a290746ee34ca

Observation 3a047d48-282c-4393-929b-712f74b3d7de · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.748945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:9ad2e42984aca4165e56dd63247d174e8378f4886561296bf681c59f28a94567

Observation 544d6ec0-1624-4948-9a40-0194af737028 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Cambrian-1: A fully open, vision-centric exploration of multimodal llms.NeurIPS

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.819030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:17cc148a545a8d724c62ec8faf19d51afd0e7b917f1ac2b35c0b7888166e8d28

Observation cc19099f-af49-4ab4-aafe-65c0ab765637 · outbound

This paper cites Vidu: Ai video generator.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vidu: Ai video generator

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.840116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:71140dadd01240c76a84d0bf7ae6b6cfada914dcced760906118fb3f89ed8538

Observation 1939d2ab-5fdc-419b-a1e5-f76610f2ee9f · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.754712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c1a31d713036d1b57b4c7388799d3b8feb67579490b7680a0ce313ec56d6ee29

Observation 51c6301b-3c18-4224-b18c-e901feadcb94 · outbound

This paper cites Exploring clip for assessing the look and feel of images.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Exploring clip for assessing the look and feel of images

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.829234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:1a410ace8ceb51532a6376f02d438e7ff3a2eca6f0a0c521426965743358d6d2

Observation 62c2ceba-f59d-49ab-8d96-4d3dc75b34fc · outbound

This paper cites Vggt: Visual geometry grounded transformer.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Vggt: Visual geometry grounded transformer

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.831507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:a6c1eb394853e951a622e8ae79dbf1569c49f82f01c9e1389011d75adb5082d9

Observation 03bec716-023f-4eae-bdb0-2b5a876067e0 · outbound

This paper cites V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation V3Det Challenge 2024 on Vast Vocabulary and Open Vocabulary Object Detection: Methods and Results

Reference 82

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.673942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c582661f66c3d5c6667b58dd3dc6cdf36e30de77605c62e27530ccd50740322d

Observation e3eddba3-60ac-422b-a8cf-c18f06e73ced · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.686633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f5715c04ef91bad95168575937e8f39b1726bdbd89c1d5f2e2b9a1d03ef806fb

Observation d8f05874-1ceb-400b-8df4-e660e8ca29c1 · outbound

This paper cites Qwen-Image Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen-Image Technical Report

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.703642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:89ccd9526c14a00b2e506ce5b6a1c214a64653fb48b46480225488f72f77aa40

Observation ca22cb71-1ee9-484d-ae73-97c5cd52e08f · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 86

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.726850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:cece68c9587b3566ddff599c4effea5442286a44b66fffb57071503fa646d0f4

Observation 71eb7b9a-4e8c-4d42-bb1e-43099e67967a · outbound

This paper cites Alvarez, Jun Gao, Sanja Fidler, Zian Wang, and Huan Ling.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Alvarez, Jun Gao, Sanja Fidler, Zian Wang, and Huan Ling

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.693583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f2c625b0e09de68a788543452058db11e570ba21b3a038aa6337f88fabbc05af

Observation de0a1c2b-5f2e-46ad-ab30-fb14621b935e · outbound

This paper cites Grok-1.5 vision preview.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Grok-1.5 vision preview

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.833384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:4167317bf7db8ffaa1ee988a9b8b1ce24120e8e41e63e8ee02ae2ceaebf11a12

Observation 9605d07a-404d-4606-8547-0ce231fe3026 · outbound

This paper cites Dreamomni2: Multimodal instruction-based editing and generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Dreamomni2: Multimodal instruction-based editing and generation

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.634595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:e5889e815bec1bbaf5432e2fed14de1d49205aa8f4b01e06c4230e75bb7ddd05

Observation eaf52597-4aa9-44d8-b578-cada8fdf0b21 · outbound

This paper cites Omnigen: Unified image generation.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Omnigen: Unified image generation

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.838163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:c69c8e70f5b0fde277c8021ce944795ceed93ee3ddfa21e9d3f2722046eedb57

Observation f35ec559-82e5-404d-9f3a-262d9319b7cb · outbound

This paper cites MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.751849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:06f73a7a10a611b9238dd4d1cbd803291ad27a0848703c75b846e8b0ea1b1066

Observation 47fafb5d-8773-4a65-a555-1a1003b673a0 · outbound

This paper cites SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation SpatialEdit: Benchmarking Fine-Grained Image Spatial Editing

Reference 92

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.764540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:e621dc68bb61db1e18a5d512ac0bdba7d2ae394507fb047a9ae464c5d92b46f6

Observation fb2ba9fb-ca4c-4815-b554-783942a8e864 · outbound

This paper cites Qwen3 Technical Report.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Qwen3 Technical Report

Reference 93

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.780141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:b1a1cf831ca01e3506603f2f52c25b25549d46ffcf139ba4c299e67d6a148bf1

Observation f7333f8b-7233-4cd0-82f0-ddfb978dc482 · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.848478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:5328f7f0471a034e4306ef3a7cbc915d376fbd2d200aca6c6c4458d1b257742e

Observation 3ce33607-ebf0-4035-a00d-70998431b80c · outbound

This paper cites Thinking in space: How multimodal large language models see, remember, and recall spaces.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Thinking in space: How multimodal large language models see, remember, and recall spaces

Reference 95

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T08:19:52.783234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:16f0255e20b1e3e36fbc49b3f7d696ac0a24ffb525bc2e85f981cecf4b7df7fa

Observation f72dfe91-3bbf-45d3-ba9c-7bd8c2e6dc9d · outbound

This paper cites MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation MMSI-Bench: A Benchmark for Multi-Image Spatial Intelligence

Reference 96

Resolution
verified exact
arxiv_id, observed 2026-05-26T02:02:26.295966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:4f03657f7e1f450f40d8dc7864caedfa6b5e72a4e02f3f0ebf5ceca41b3d6707

Observation cbfcce05-18cd-497b-85f1-4b8ddc6a3cff · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 97

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.696934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:046302c3dbb5db8d35379c314a12796f70698a5a99feee50aceab2e21a784406

Observation 38df0576-8768-4231-942f-888ddd89cf95 · outbound

This paper cites Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Seeing from Another Perspective: Evaluating Multi-View Understanding in MLLMs

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:19:52.798498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:2edc100c5dfb4d9a1baeb1dd7ee25269872266535f60cdff8485a2993e9d0ec1

Observation 9f6c08b5-40e4-470d-8125-f9384766fe12 · outbound

This paper cites Scannet++: A high-fidelity dataset of 3d indoor scenes.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Scannet++: A high-fidelity dataset of 3d indoor scenes

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.878214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:f1c7c1c5ff8de3f3d86e38b3a14ed47ef95660cdb58a75dfaecffe50284b0f6d

Observation 1df1c87b-61e2-4aee-88fa-de82558a3546 · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Anyedit: Mastering unified high-quality image editing for any idea

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.827328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:4de937370a1716973a839b4ecf481e4e87856f3e473865970e3f91f6c4f83d94

Observation 978e6a56-9ef4-4171-8265-f8268c3936df · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.Advancesin Neural Information Processing Systems, 36:31428–31449.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Magicbrush: A manually annotated dataset for instruction-guided image editing.Advancesin Neural Information Processing Systems, 36:31428–31449

Reference 101

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.821009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:74dd32d57551fc4d400992f4c5a49249b29dc4ce72a4775cfb85fa6f280e242d

Observation ab95a6e8-f54d-41a3-afe0-f89faf92c57a · outbound

This paper cites Bamboo: Building mega-scale vision dataset continually with human–machine synergy.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation Bamboo: Building mega-scale vision dataset continually with human–machine synergy

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-05-21T08:24:52.825143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:cb019d0353b02a80ac07ebbfb7264356a7855252deea5b9a18a5e6163ebfba8c

Observation 3e43ecc3-b147-4a01-9e6f-b3443548d234 · outbound

This paper cites In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer.

JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation In-Context Edit: Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion Transformer

Reference 103

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:19:52.774390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-05-21T08:15:58.020894Z digest=sha256:fc45905988a9c968f7676f67fa11a0735f770765e8deefea2071cbec85acf8c5

Pith citing papers

Observation 305321f0-c390-4fac-b274-50ddbeb32e94 · inbound

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models cites this paper.

Embodied3DBench: Benchmarking Low-Level Embodied Spatial Intelligence of Vision Language Models JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-29T13:13:27.222737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-29T13:10:54.222870Z digest=sha256:6461c18537e7620ddabdfcc65d496f3b826ad587a5161a2712b8629af20c2a7b

Observation 91e3a084-0cfb-4708-aa47-020927d0415e · inbound

Qwen-Image-Flash: Beyond Objective Design cites this paper.

Qwen-Image-Flash: Beyond Objective Design JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T02:26:26.222731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-06-28T11:00:04.775189Z digest=sha256:9a8ccbc19924c7340ef38b5de232e61b050f22e47b34250da9c0f6121747fd53

Observation 372dea67-29d7-4177-82ad-c6ac9e1b8d89 · inbound

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation cites this paper.

DetailAnywhere: Fashion Detail Generation via Cross-Modal Feature Alignment Distillation JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 77

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:58:37.268189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=pdf_text observed=2026-07-03T15:56:38.037304Z digest=sha256:4cdb49ab1a1aa642a83f7255fc2a344c77616150d6bb5bfd4bc51ca640302270

Observation dd050f85-7c63-4513-bdc2-6e6af8183190 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:03.146944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:03.146944Z digest=sha256:a3a93fcf1e68ec908653f1aa4dff0c731a0aedc81f935487f7cffb1e388410ce

Observation d96657bd-7282-46d6-9cc4-98e0a09575bc · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:03.262803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:03.262803Z digest=sha256:f974e37aea47e20e12a07d88c2ba0f05120a66964a4a50b068f8275aa3485a3a

Observation c1a5d00b-283e-4073-964a-8a37a63f970c · inbound

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On cites this paper.

Oxygen-TryOn: Fashion-Native Foundation Model for Any-item Virtual Try-On JoyAI-Image: Awaking Spatial Intelligence in Unified Multimodal Understanding and Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T07:07:12.634593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:07:12.634593Z digest=sha256:7a9cc7b718992fb053fe3618da89c31fb86720dcea59652b063af2e9e4afb85a