Pith. sign in

Paper Citation Record · LEDGER

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation

As of 10 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 0 inbound Pith citation observations for arXiv:2606.04264.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.04264 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T10:23:43.501656Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

69 of 69 outbound references displayed

  • verified exact9
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch43

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ae7a97a6-e716-4f44-bf1f-79e6625f9a38 · outbound

This paper cites Qwen2.5-VL Technical Report.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Qwen2.5-VL Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.566595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:45bd2d57427c5920b118402bd36d4def91ada0acc9bd0777faf91c25e59e62d7

Observation f3e6aa21-e713-40fd-beb6-3c0cfb76cf7d · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Con- ference.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Proceedings of the Computer Vision and Pattern Recognition Con- ference

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:998b5ae38e5ee922829d8f562546647061d7d564c6b6cb1453dbf9628b903d40

Observation 9b1deed4-1123-4032-8db7-38b6d69b0b3b · outbound

This paper cites Show, don’t tell: Morphing latent reasoning into image generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Show, don’t tell: Morphing latent reasoning into image generation

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.548371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:80110f5f9cdc12ed6e94733f968ebd671d6395264ae91dee88ab378aff5ada20

Observation c528c30b-c05b-438e-b637-77a36c5a076a · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.539468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:6155bbb336ea3815911ab225c77150c525c3e09a41ed64065cfe7a7686fec11e

Observation f0e0db20-42fd-4e0c-9467-3b005188e681 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.559437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:252cd5c26de6f3bc79bd25dff3a05acf87c8957432a08161bced70d088c6b154

Observation bfe18290-46dd-4369-b218-fcd1ebcb5253 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.571000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:26aa26d5f4b46cdd43cb1113ccdf8747f89a64685575d4a6e58ec94dc5fd0328

Observation 70a67ce8-bb71-4b32-9176-df6bd433ff6d · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Emerging Properties in Unified Multimodal Pretraining

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.567184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:5f988435726611264378cf85e91c50c9cdab97c42170201c72d84f41789ab98d

Observation 7c6c606e-68b6-4ccc-ba07-b40f4442edbf · outbound

This paper cites Advances in neural information processing systems34, 8780–8794 (2021).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Advances in neural information processing systems34, 8780–8794 (2021)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:481b7d15d59d2d8149ce4573bb04c46987bc16e0339f47bb6a74b8956a4e49be

Observation 094f7303-a6cd-4d0c-b58c-f37abd8c2d9c · outbound

This paper cites DreamLLM: Synergistic Multimodal Comprehension and Creation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation DreamLLM: Synergistic Multimodal Comprehension and Creation

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.585749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:87d40f68d17d4bda7f221a30a15294e3034e20eebb5409d72a1adb6f461facb0

Observation 47ed9e8d-6ae9-486a-b062-8ba11c7620fc · outbound

This paper cites GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation GoT-R1: Unleashing Reasoning Capability of MLLM for Visual Generation with Reinforcement Learning

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.575568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:6d3babe1b740f4348c805516febee36f151a4e17f9ae982ccd0d1544fb418a01

Observation 5877e331-b0f2-40e5-b34f-ac22629605fe · outbound

This paper cites In: Forty-first international confer- ence on machine learning (2024).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Forty-first international confer- ence on machine learning (2024)

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:0c8535808adf2f2e8fc25dcbbcd3e97a80967723ecfb5f2a852df1c7bb5f0015

Observation 80f432bc-4c07-4351-a531-359f6798c77f · outbound

This paper cites Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Lumina-T2X: Transforming Text into Any Modality, Resolution, and Duration via Flow-based Large Diffusion Transformers

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.553020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:eef99543c2a37b49b49066736b11bc8bc82e5d1a0c545a35e7d3a3abe883b8e7

Observation 6f90cd72-d9aa-40b4-b4db-5dc51e6b16e4 · outbound

This paper cites AdaWorld: Learning Adaptable World Models with Latent Actions.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation AdaWorld: Learning Adaptable World Models with Latent Actions

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.534534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:10f1bcbfb484d3bc1c0eb1e62b224e31199fcb54940c304989ffa0b5373b4cc1

Observation 6a7cfd64-7b34-49e5-961d-1f152a81f8c5 · outbound

This paper cites Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.582230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:c36d2c7680e9d212bdad76168dcbf5e9820e432ced12dbd55d7d6a2cdcc0f616

Observation 0986b970-58e0-4a48-b9d2-5ef374d3651d · outbound

This paper cites Thinking-while- generating: Interleaving textual reasoning throughout vi- sual generation.arXiv preprint arXiv:2511.16671, 2025a.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Thinking-while- generating: Interleaving textual reasoning throughout vi- sual generation.arXiv preprint arXiv:2511.16671, 2025a

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.507525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:db2f4de5cff7c241950a2a8e808b45d95fefa6c608c032f51a4208b8c6714970

Observation aafb7341-2c9b-4aa3-a2c7-b6e392123764 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.520773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:790bd9f028d99761d6110ab542a76decdea125cc2b78ba208884813567428114

Observation b664f2f0-9294-4396-9ddc-0f0947df1b26 · outbound

This paper cites World Models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation World Models

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.592911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:be9b8747cad2f85dbbacab5828fa0d7d94ffc60ddf43cae3c6219b06cf4cf499

Observation 32a4cfd5-cd0e-4da7-8b66-8cbf9793f918 · outbound

This paper cites Dream to Control: Learning Behaviors by Latent Imagination.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Dream to Control: Learning Behaviors by Latent Imagination

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.542661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:a6cbce822939f8048be159eb69dd1c9b3dd52858a78df56a8762eb116228add5

Observation 0500a625-4ba1-4266-85ec-7cbd6983dfff · outbound

This paper cites Unicorn: Towards self-improving unified multimodal models through self- generated supervision.arXiv preprint arXiv:2601.03193.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Unicorn: Towards self-improving unified multimodal models through self- generated supervision.arXiv preprint arXiv:2601.03193

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.441036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:ab48b77aaa986eee3a7151e2d295685f033fc2be29b1cdee8d50802bc6e6a2b5

Observation bbdb05bf-ec04-4638-a7a3-db54f6b51430 · outbound

This paper cites In: Advances in Neural Information Processing Systems (2020).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Advances in Neural Information Processing Systems (2020)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:3e5f6beadf5a0b96bf9fcd36ce91d2a1b978ddaa8321e70864ff136f6830c1cc

Observation a6645927-226f-4371-ae94-11dea5f5c1e9 · outbound

This paper cites Classifier-Free Diffusion Guidance.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Classifier-Free Diffusion Guidance

Reference 22

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.533665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:4673d5a1c600a70c3d7f336650c2692080af903c2c0f01d3174a6472c9b32d51

Observation 8518f4d6-fe6e-4352-b729-87ace7b4ed5b · outbound

This paper cites IEEE Robotics and Automation Letters5(2), 3019–3026 (2020).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation IEEE Robotics and Automation Letters5(2), 3019–3026 (2020)

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:0c0369b5b47276dc5c06658557e8e1642a59b93109fa404450d538abfcb84081

Observation 4e41e694-401b-4ea4-a1c0-b39ec79e72f4 · outbound

This paper cites T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoT

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.453307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:8364e7fb985b12f21c804f76952ba8f5dbe3cae8fd6fadcea70462c01aac5254

Observation 364bed1f-510d-476b-aa6f-089093e51c26 · outbound

This paper cites Draco: Draft as cot for text-to-image preview and rare concept generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Draco: Draft as cot for text-to-image preview and rare concept generation

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.597193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:a60fcfa47cfb036a2674dc45ac5310ea1cbe32b4e494b88c6de8eb833787b2f6

Observation 1517d4da-aaf5-4c34-9af0-062175169ba4 · outbound

This paper cites Co-reinforcement learning for unified multimodal understanding and generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Co-reinforcement learning for unified multimodal understanding and generation

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.588932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:ebb8f533b138ce16228cc4d57b5bed0184baea88fc6436dd7cf1ac21c9495a9b

Observation 6b582f28-6385-49bf-be00-bf52174212e9 · outbound

This paper cites Advances in neural information process- ing systems35, 26565–26577 (2022).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Advances in neural information process- ing systems35, 26565–26577 (2022)

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:c0366754a3b0543a1333395885a7fe57c0272ac9f4dcf0738609a35210aea1c8

Observation 5cc6122f-c3c7-405e-b5d5-c6111f103c9d · outbound

This paper cites Advances in neural information processing systems34, 21696–21707 (2021).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Advances in neural information processing systems34, 21696–21707 (2021)

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:4a3d1ed9af74a31ef33be95d7c67a7d5e2e11416196e24d1e1139df4a4574b6f

Observation 0a89f193-74e6-4fb1-bbba-d4a4be71e2d7 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.480653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:f4b5df8c641f9219d6688217a6080245d90678e9f24a87317608fb0e73b02dcc

Observation 60030a6d-a422-41b1-aeec-878c0af925a0 · outbound

This paper cites Causal World Modeling for Robot Control.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Causal World Modeling for Robot Control

Reference 30

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.604172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:98945e1508a37c4e852336e03b125e496bec9588fcc0c33d7df6be047daae66e

Observation f68e8e09-fdd2-4091-9d33-9828d4d24719 · outbound

This paper cites Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Genie Envisioner: A Unified World Foundation Platform for Robotic Manipulation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.499211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:d71ff1ffadd9f380140dcddececa81a40fb1029676bb588585eb38e71ca91803

Observation 92ae4fb1-18b5-4653-b5c0-32b1946bd869 · outbound

This paper cites In: NeurIPS (2023).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: NeurIPS (2023)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:eb93f3f903b248faf6bb8be68c83d4befc312c6980b2f9e42dcb850c7839b67d

Observation 10d35a61-b93b-442a-a6f5-cf6ebce8b372 · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.529581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:21d6b7e6a88efe2454ccc70de49e73b58f0776900ad168dccc9507ab3b14284c

Observation b4913942-e06f-4c3d-9ffd-70df057d8e48 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.475775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:bc10fcce5f69bdc25fb57bca6becb6e299344030a30b92947f7c0b12756f4eeb

Observation a044656e-8c97-4103-a713-dfad79292136 · outbound

This paper cites In: International conference on machine learning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: International conference on machine learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:cffb7d874dd37130268ad3e97c88111eec081d70f05a803db2e0b97c8d3e5edd

Observation ea36146c-2f6d-451b-85d9-e64082d4ca1b · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 36

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.457737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:7db78a6ea87f855206a110f598411aa5c0417018bae3f62b2f82d97c5b7895ce

Observation 3a6354a8-9b1c-4ca2-9942-0557b5ae1097 · outbound

This paper cites arXiv preprint arXiv:2510.07313 (2025).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation arXiv preprint arXiv:2510.07313 (2025)

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.468681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:80d65cb56b615049256c7582bfd3ad7ed387faec192d30a641fe62e96a96c7f0

Observation c67e72b4-edbf-429d-bd8f-80996dc335f3 · outbound

This paper cites Uni-cot: Towards unified chain-of-thought reasoning across text and vision.arXiv preprint arXiv:2508.05606.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Uni-cot: Towards unified chain-of-thought reasoning across text and vision.arXiv preprint arXiv:2508.05606

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.511892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:4449c52041825d07dac4c425beb1d8eec05c760bf08f2cb1d2a95d42de929fc3

Observation 465ec9fe-d271-4ec3-ace3-945de1659d2c · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 39

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.503458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:ef955293381209efad1c794a2c49facab75f019f659d2c818ed4527395044607

Observation 06d7010d-177b-4bca-a8e2-214ec9e59a7b · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:af57bb87dd8181ad4ffcbd9dd4bd756d59ebe9ba3b3e9b9dc209ce2502bb41f5

Observation 932508e1-a945-4d8f-9371-108ff7965203 · outbound

This paper cites Advances in neural information processing systems35, 36479– 36494 (2022).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Advances in neural information processing systems35, 36479– 36494 (2022)

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:0798412b6c61c97bd44c3c0f100dfc92a765b493b9c94d92e0eaa640957b8457

Observation ae49b2af-536e-4ec7-8d72-cf2a1aabbb55 · outbound

This paper cites Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

Reference 42

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.408322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:a59699a60b4e53b73527184a7eea4bb0491e516988cc70df03a76e972fbf2959

Observation c7fa0d79-cfc7-41fe-b50d-1bf061b75162 · outbound

This paper cites In: Proceedings of the 32nd International Conference on Machine Learning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Proceedings of the 32nd International Conference on Machine Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:7e5e6377eca309d46051817d7b4b8a2f2175737966a730034a6d84525320d1e7

Observation f92ea737-4087-4416-8475-3a3a4c6e3d01 · outbound

This paper cites Advances in neural information processing systems34, 1415–1428 (2021) UniCanvas: Diffusion-base Unified Model for Text-in-Image Joint Generation 27.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Advances in neural information processing systems34, 1415–1428 (2021) UniCanvas: Diffusion-base Unified Model for Text-in-Image Joint Generation 27

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:6d2ef2a92c9058966c6dc744ba2ac0e9600ada99273d94601add0f570147df21

Observation 73f7e765-0833-4779-a4ca-35536a43a67b · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Score-Based Generative Modeling through Stochastic Differential Equations

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.529458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:7d721b4ff8b822003d39035dadc68c78dbaad4d63f087c97127e942e941dc5a0

Observation 3b2abf4f-2aa1-47f9-9a11-80e9d55834c1 · outbound

This paper cites Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

Reference 46

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.600947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:21e59946a256f18a3c5455f5259413f835be715b1e4c065e1e7c40f3eb5b34fc

Observation 39ebde8e-370e-4b84-b3f5-9509fdb7be11 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:29.498536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:69b9456e69540878517e8a95c6b415c35fcb53b097e1687c0403dc762c26d2b7

Observation 19dadc12-d78e-448a-8d84-d4f116ea6a92 · outbound

This paper cites Next-Latent Prediction Transformers Learn Compact World Models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Next-Latent Prediction Transformers Learn Compact World Models

Reference 48

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.434904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:504de0a9ee51f40fed5c5f3f31b69740cf6eb48d68f18a1df5e00fd4734113d4

Observation 3bff77f8-fabd-4154-9dd9-2c0556a56736 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.515827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:f02fa63c80827d56dd00ded353a86e9f726006e3830d20f014b3d923641ab522

Observation 560668b8-fa32-46fb-9460-09c127905331 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 50

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.520043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:b7a925f1575298a34f618f74913cdf64ad18ce1ca1fa957a072b599fac5a848b

Observation 63aaf219-5335-4093-900a-5a34697599fc · outbound

This paper cites Promptrl: Prompt matters in rl for flow-based image generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Promptrl: Prompt matters in rl for flow-based image generation

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.431227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:38c543800a72d5f4d05c62a06ec3d253b4b5b650679b6f55569abb57bffba0a8

Observation 2d614183-073a-40d1-b8dc-a1449344745d · outbound

This paper cites FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation FUDOKI: Discrete Flow-based Unified Understanding and Generation via Kinetic-Optimal Velocities

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.415805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:ec106fe9baf6e27614e53447b9bbed62cfe9cd38e323e5578a881b1f2ec6ec71

Observation 3f2a96d2-7e12-4b69-88e6-03684f390914 · outbound

This paper cites OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.555038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:1fd7bf8b50e0d3fec54e6f1404adabd2e3f592b30063ff1f265389614de42f39

Observation d8418e6d-b70c-4c8c-8d8d-ce2017c5df22 · outbound

This paper cites Visual generation unlocks human-like reasoning through multimodal world models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Visual generation unlocks human-like reasoning through multimodal world models

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.404589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:8dba358117cfaedf80e0333f7178d350687ab239843d8667c5c7f50b9d091bfc

Observation 18f693bf-10c3-470e-ac10-01cef63244f4 · outbound

This paper cites VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation VSP: Assessing the dual challenges of perception and reasoning in spatial planning tasks for VLMs

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.516632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:bfb159fbf2ed954e1fb43ae0554fb767fd8b41944cdd9af56270a73d7425ef48

Observation 111f96ee-2a70-4f09-b5a9-511d8a3cb994 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 56

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.562678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:f6b59136904413e45b3ded77a2eb106ef22017a47b44bc190c815d2a1fb23221

Observation a6585bc5-4f90-46ac-91cc-5b059183f3b2 · outbound

This paper cites MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MindOmni: Unleashing Reasoning Generation in Vision Language Models with RGPO

Reference 57

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.442237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:19ac5abf93cf6b6c480145f1d77b4450650dc8d3c7859309ea107effda38f607

Observation 83fb98d1-c4f9-427b-b126-9c862489be1a · outbound

This paper cites In: The Thirteenth International Conference on Learning Representations (2025).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: The Thirteenth International Conference on Learning Representations (2025)

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:c4e2a8ee9afce4b89237af698b9000cf01bf360c5898822a08c387bd3d32a219

Observation 39258149-5fd1-4bf4-9599-7e73c78c7313 · outbound

This paper cites Visual planning: Let’s think only with images.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Visual planning: Let’s think only with images

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.570311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:0d94fb4959633a4174608031961059a0a2ba2af94eec0f5be8702011f1034a8f

Observation aa8649db-da31-4f93-b55e-85eee930f019 · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation MMaDA: Multimodal Large Diffusion Language Models

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.573608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:c6d7c05b0eb29cee3a4749ef0ad7068e69a96c5d13751aecdd8d42c3eccf66d5

Observation 3cb7cb13-78ad-4b68-8ec5-f6e9d87dc1f3 · outbound

This paper cites Mindjourney: Test-time scaling with world models for spatial reasoning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Mindjourney: Test-time scaling with world models for spatial reasoning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.546873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:34985cfa63c026dcba60b1b55d273a591fed8721f46ab004bad89fdda5112f46

Observation 63ddab03-9136-4c40-b3b6-1604641457a1 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.449658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:b4f258fd445554e1ee2ea456cb596440e929aaf40489eb409e73b85fbac0888a

Observation e6ce7c12-2569-4435-a292-498b1b211c7a · outbound

This paper cites World Action Models are Zero-shot Policies.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation World Action Models are Zero-shot Policies

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T03:06:29.466690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:1c5bf3b1eec5d55320b3ab6c02f9e63965cd6cd4eff34212fcd64ccc5b55073a

Observation 1e0ffade-cdc0-4094-8001-02d4c6d40dc7 · outbound

This paper cites ReasonEdit: Towards reasoning-enhanced image editing models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation ReasonEdit: Towards reasoning-enhanced image editing models

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.550865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:dbedd88418956ef1e5dfb66f532217bd2f839a4fa2f92d8bf73372b7505edfb0

Observation 5e1ea126-ba6b-44f8-9c9b-0ccb131017ba · outbound

This paper cites In: Proceedings of the 33rd ACM International Conference on Mul- timedia.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: Proceedings of the 33rd ACM International Conference on Mul- timedia

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:e6e1d4582bd4b5fb3f66adbbbe8c57e0fe24b0cf07232ec99f2c50de24045965

Observation 4f3707d0-f92b-45ca-a21c-50531d9acb59 · outbound

This paper cites In: CVPR (2018).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: CVPR (2018)

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:c3641f712412308bc57c4c2f46422d32156d0fe7f404120e525eeddbe5390139

Observation 975754a4-f591-4dba-a68c-5dd9242ba140 · outbound

This paper cites Foreact: Steering your vla with efficient visual foresight planning.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation Foreact: Steering your vla with efficient visual foresight planning

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T03:06:29.445867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:67a40e501fc3e6149a31392030f3356f01882727bf4e48561bbba728791dda28

Observation fc6502bd-1ccb-4aa6-b95c-58b9b48bac7a · outbound

This paper cites 3D-VLA: A 3D Vision-Language-Action Generative World Model.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation 3D-VLA: A 3D Vision-Language-Action Generative World Model

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T03:06:29.470876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:4c3add437ea703109949f1667d5989c0558d690523608b6346700d99b6e4e103

Observation 9d313dfa-86a2-4787-ba74-5da90799e4b6 · outbound

This paper cites TesserAct: Learning 4D Embodied World Models.

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation TesserAct: Learning 4D Embodied World Models

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-02T03:06:29.578809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:de5f0c1641cc0c3b88388ea87732fb8dcf1fad554ab7b258166a3b671d1cc8e4

Observation 73a92f2d-182f-4321-a9f6-5505505a230a · outbound

This paper cites In: The Thirteenth Inter- national Conference on Learning Representations (2025).

UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation In: The Thirteenth Inter- national Conference on Learning Representations (2025)

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-28T10:23:43.501656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T10:23:43.501656Z digest=sha256:af04cc64713597ee966e3f306f162c9f2391e34d686c02958b3d0241245637ad

Pith citing papers

No inbound Pith citation observations are available.