Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:18.590800Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 26 inbound Pith citation observations for arXiv:2505.23661.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:18.590800Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:37:16.161813Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T16:39:58.237442Z
73 of 73 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6035d8ec-a76e-4ea3-b87f-954d67ad2cc0 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4a14748-5626-46a2-b2f0-d7828f34232f · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Visual instruction tuning, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 578b26f6-e097-41ca-9f47-f8e52869ffde · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7054ed70-b2e0-4bd7-8a9a-eef25feca87a · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1ff2ac-df0c-4395-b94e-d6748d81253f · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bf26c07-2194-4332-bacf-f1af516967b0 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd256ac3-fec7-4db9-a8f9-e796a66a9bb7 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48873b2e-ed5e-451d-ab14-b7be49289677 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7932a8d-5df6-4975-9d6e-a08bf667421e · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f3ea260-958c-40a7-ba07-2f6de7deef77 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SDXL: Improving latent diffusion models for high-resolution image synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc33ce15-5a5a-40cb-91cd-f8ffe9b554a6 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ee0b66b-1621-422d-baf0-e6c2443cb82b · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Zero-shot text-to-image generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ea54124-5ae1-4e4f-a7df-a3851213fbfe · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f5966fa-82c8-4c95-b511-30fee3c82857 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improving image generation with better captions.Computer Science
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ed071f9-acf5-477a-852a-ccd165990bc6 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Attention is all you need.Advances in neural information processing systems, 30, 2017
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c87388e8-1a6a-44c0-a5e4-6a5d94ef3bc6 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation GPT-4o System Card
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34ab3c70-925c-4147-91a7-c49f487e21fa · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc3aa865-95bc-4299-a5b2-ae5b1f852c80 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4ff7499-3701-4273-bee7-aaf6c91ab774 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation feee7fbc-8865-4c9e-a27c-61780c8fae68 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e20cbcff-517f-4cda-a019-6c68bd4f3b08 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7590a71-3091-42fb-ac32-7b3bb6302515 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9dea184-515f-4a16-8c6f-2a99ab13ca1e · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c9ce947-b55b-4545-be2e-997b370492f3 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 189049c2-f882-4d64-b9bb-2f825baa2e5f · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10b76f1f-2460-40a6-a315-cdc7ee7d006d · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fcac129-639a-4db5-b51a-66727b97a7e1 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Transfer between modalities with metaqueries, 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c947697-6a78-4fde-b22c-7779b8e343b1 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41ef8a11-9e40-44c8-b07d-8ea86291d395 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision, 2021
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d49b3c90-0a92-413f-a3f3-805b76b068b4 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre- training
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaefd3ea-0f7b-4612-aeaa-352193c3059c · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5573576b-e62c-4c13-8f53-a4148241707c · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76af517e-7ec2-40fd-977e-99a915ce80c3 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92061989-3c66-4e66-9699-a0e2baa14917 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d89353-e2df-46d1-9085-6e6aed608311 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39b5e017-0143-4b75-9849-cd0125d4f9fd · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07610958-a4b0-4e79-be63-b8bf11b1e917 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scalable diffusion models with transformers
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2077f18a-effd-48eb-be04-092b0ee7208c · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec1f466d-522f-403c-9f96-6012750e4d9b · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ab9c62b-e79b-4b6b-97f6-a13c8b0f88f7 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation U-net: Convolutional networks for biomedical image segmentation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c60d4210-7d90-44e4-b06c-ea52c35bfb5e · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ea1f9f1-6990-4fd9-b2fa-c90bcf99f027 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4dfaf339-61e1-49d1-8b58-eda1097cfc44 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux, 2024
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19350a6c-032d-4ce6-8ec1-3fb0e20ebb8e · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f46da92d-f9b5-4608-99af-db1e4c6520d2 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76282589-2e37-47ca-b250-694a41e03f99 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38a85164-1543-46b7-b883-2d733592397b · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8d77a9a-3258-4780-8bc4-d4aadae3a339 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6378ad2-21fc-4c2c-939a-1eaa4ae8a8dd · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation F-LMM: Grounding Frozen Large Multimodal Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e059d4e-8c73-46df-a29c-1c8a4bb510f2 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling Laws for Native Multimodal Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13c031c1-c6c1-43ea-8e5c-f94f3a17d559 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f667d273-2f2e-47ac-892a-b5ff3481590c · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ea9665-0bbe-4b92-b2dc-81d09ba090cf · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef601de-0ec9-4569-b8aa-0709d22d821b · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d648e502-14c5-431f-815f-920dd602f54d · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models, 2022
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0762f47d-604d-422f-92ef-14c98bdbfb6a · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d7ba6e-4e7f-484b-8866-57a44e081f07 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b128eb78-e9c5-4581-a077-ba2022a1ad64 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flow-grpo: Training flow matching models via online rl, 2025
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdcc7b72-102b-4a33-a9df-31ce5984a324 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15e45610-e6f7-481a-aac5-478ccd9f27d4 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62f640ca-edc3-42c1-af76-465ed47dc62d · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010c2984-c358-4ce9-b2c1-2177b04ea6d5 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation text-to-image-2M: A high-quality, diverse text–image training dataset
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 538e3db9-6637-4525-a4cb-c950e8931524 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb5006ae-20b2-4e04-bedc-c6914bae2c8f · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Megalith-10M: A dataset of 10 million public-domain photographs
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c1e9cfc8-45c1-4728-bf5b-8f6682335c2e · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation RedCaps: Web-curated image–text data created by the people, for the people
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 44765e00-e05b-477d-8811-bd652baf00a5 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c6db8f2-4aab-47be-90c5-330ac5c78f80 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eea42e66-daf8-4d68-9b58-bb75b8c65925 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47103e9b-bdb7-42c9-94ba-44a69aa07525 · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · outbound
OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b93d0b8-5e55-41ec-8c49-81e8ff292d4f · inbound
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2440898c-8e79-4572-8e74-09e812db5517 · inbound
Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6e71b26-80af-437d-8272-7790425ea856 · inbound
Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2dd93c32-9928-4ffc-a9d7-b37b27e91874 · inbound
Reconstruction Alignment Improves Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1120e7a6-8acc-4195-9c58-944c4448e8fa · inbound
Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21aa7041-6a31-430c-b826-375928f15e2f · inbound
InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c3b208eb-af74-4035-9ed4-2c9adbb461d9 · inbound
Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e93ea3-9f3f-415a-aa55-0158c10054e5 · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b0f89c4-257c-4a73-a01f-98ebefdd4451 · inbound
TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c96aa05d-13c8-4ef0-88af-bcbcf5319edf · inbound
Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8b082f0-af4d-4067-a289-d77dd57f1849 · inbound
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 519af2ee-a3a0-4291-83e1-c444c03131fb · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · inbound
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ac2fb108-0f9b-431b-8e5c-909ddd734640 · inbound
MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 146
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 333d691b-d2e0-4bd4-9cd5-d6ba23abcb42 · inbound
Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 008a8304-022a-4fbb-9b37-34e5fbe1fa6f · inbound
UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4473075d-f5ca-4df6-8206-6ce5471a2566 · inbound
LatentUMM: Dual Latent Alignment for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6a177bd3-2801-445e-91e8-250c9704f213 · inbound
Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a318c7ae-ef0e-480b-9093-0fd3fadac828 · inbound
Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 56505dc3-693d-40eb-bc29-692615926173 · inbound
HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8243c92a-4a9d-47b3-a384-fee137bb55cc · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e827ec3-e9be-4230-b11a-982702ef9b13 · inbound
SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 962dcbc7-89cc-4983-8fef-605c13eafe22 · inbound
IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4335c52f-26bc-460c-b559-203232a2b955 · inbound
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ae6621d-7e99-4606-8967-b66ff2bf033f · inbound
IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d74a6e2-0343-4150-a5f2-213031a8a0d0 · inbound
Test-Time Curriculum for Open-Set AIGC Detection OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.