Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:55.930004Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2505.14682.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:55.930004Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-11T22:05:33.892978Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T21:37:25.273563Z
95 of 95 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 52f83fea-7511-4ab7-a0ec-43ee187f46b5 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ccdf867-2c99-446c-9a6c-d8ebb6f0135a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Flamingo: a visual language model for few-shot learning.NeurIPS, 2022
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a29d6707-0930-4c4f-91ba-124d843d6482 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 723682f7-77e5-40ed-95d9-d0776ccbd48f · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Improving image generation with better captions
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dd1ae7a-d72b-4dc4-85c2-b208561c3573 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Maskgit: Masked generative image transformer
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41889b5a-3469-4d56-bcc6-237abb7e73b1 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd00a909-2791-402a-9890-700fe45a6f9e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sets: Leveraging self-verification and self-correction for improved test-time scaling
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b20ee523-5bdb-4ed2-b2b0-729660110fbe · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb32b126-d539-4364-8e45-ee9b1ffd6482 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation DeepSeek-V3 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a7a45d-90ed-44dd-8235-3e0642eced0b · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c67c16d5-3ab7-46a2-917c-ff48f0d0e502 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3859a696-5813-410b-8b04-c23ae7268c3f · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb2098e1-9e53-41ce-a64b-85e74d3182cd · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26ad3188-5404-46a4-89a1-afad38099a2e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5efc5f2-9304-4718-aec1-1a86e142a83d · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Geneval: An object-focused frame- work for evaluating text-to-image alignment
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2e84395-3dc7-48de-beab-b1a69a24e420 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation https://x.ai/news/grok-1.5v, 2024
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 783a8d68-ade3-4a30-ad42-6f0486bb4721 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fc057b5-f3db-4e60-b51b-77bf28e69907 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9ae4cd01-79ad-4ead-9d2a-ada7e66c244a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53f55328-9d90-4735-b33c-25db97047411 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c8de696-4634-4b05-8551-d6ca78412285 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02d160ea-3ef4-4150-b529-792935c7a446 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Efficient Test-Time Scaling via Self-Calibration
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a48737-5dac-4fcc-af25-5e4dacd87703 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation T2i-compbench: A com- prehensive benchmark for open-world compositional text-to-image generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4168fd3d-96c2-4cf2-a42f-8dcaab03c391 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94a9d7f0-ab74-4b31-8f27-4d6f0e347e4a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback.NeurIPS, 2024
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90b0f1c6-29d5-4497-9314-3dc8be323414 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation text-to-image-2m: A high-quality, diverse text-to-image training dataset
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31234611-dab3-4a1d-a3d0-a54352ddb84b · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88929cfb-eeee-407d-86d5-fbb88bd7afc1 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation A diagram is worth a dozen images
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982002e0-8e72-4afb-9316-9fb10a7b138f · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Segment anything
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d56196af-47df-4c75-96f1-e351b73c2c68 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d629b1ae-9e12-4c2b-84e2-de35dadc566c · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbd58054-26e3-43d4-9727-aa8e40f15f80 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f4078a-e59a-4243-a4d3-993af8aa5195 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Evaluating object hallucination in large vision-language models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bbb5788-a59c-4612-a250-14e6e2bce774 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca4ad6c1-b257-4436-a767-6af5ff7bb5d2 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45e4b063-c515-41d7-a93e-82ce1dfa65cd · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30077521-d764-4357-aae6-cbc02cabcb78 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Visual instruction tuning.NeurIPS, 2023
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6368733-9fef-477b-b20f-a6c9b9c6ef7a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation NVILA: Efficient Frontier Visual Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e584414c-19b1-4a33-9fa2-6928c587e898 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a974054f-a8f6-4b30-9fba-5ae051fc1008 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3a559694-8a17-472c-8603-2ba29b57c09e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0d2967d-0f9c-46f5-8a8d-af0fee297e29 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9e736e6-6583-4c76-9921-b5aa0b8a23bf · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mm1: methods, analysis and insights from multimodal llm pre-training
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67ee5d29-e318-4817-8448-150ef028d871 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Finite Scalar Quantization: VQ-VAE Made Simple
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c6223bf-0d29-4da6-9e12-a8d5af6071e9 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation 4M: Massively multimodal masked modeling
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3a0b26d5-444a-44ca-8cb2-4ad988125365 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Gpt-4o, 2024
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ed0b6e-5140-482c-bf39-a60864f2e53a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 457042c4-d2c9-4fdd-a391-87a1a1d0f3a2 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sdxl: Improving latent diffusion models for high-resolution image synthesis
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 078059fc-f144-40af-b551-fc62dd3a4209 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f242a8cd-9bbc-4ae8-9eb6-7544a417f009 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6a5ed8e-c7f1-4399-8062-165a553d84e0 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Direct preference optimization: Your language model is secretly a reward model
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e76550e-fba5-4a35-b15a-a9b53de0db40 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34142080-8a96-4211-8d6d-048ac101e8a3 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ImageNet-21K Pretraining for the Masses
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37da172c-6c94-4ad1-85e4-c40fd9e49d2e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b34b7e8a-a3e0-42ba-9eb0-12d860fa1324 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Journeydb: A benchmark for generative image understanding
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 20e719e6-c461-449c-b2ac-866aa70f6eff · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Aligning Large Multimodal Models with Factually Augmented RLHF
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83b385a8-a10e-46ef-ab1a-eea6ce7f4a7e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd113272-fccd-4aa4-a22e-de30ec086e59 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7179324-bc70-4faf-99fe-2111eb1aa559 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 4bbe62b5-fde3-4c19-b30d-e018fee5bd08 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afa9970b-f0fa-494c-bb45-39a1cf410247 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad0c2970-0cfd-4b2b-abe8-f93d345c7370 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 64
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 846badcb-5f5b-47e3-8b07-2746accc0d45 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Neural discrete representation learning.NeurIPS, 2017
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7126fe6-580b-4126-8f5d-7bf49f15da5a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dd2d8fc-0dd2-4746-9008-f3120daa65f6 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL
Reference 67
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 797fd783-9290-4779-a178-3ce9d99374b9 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e50f3828-5425-42c2-acaf-68bb2490f3fa · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1b25499-ef54-4575-90e3-e45f4e117118 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VisualPRM: An Effective Process Reward Model for Multimodal Reasoning
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea95bff2-5f5e-4351-8b28-03f390050f15 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cc76da2-e4ec-43f6-9233-39a53873b3f8 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mint: Multi-modal chain of thought in unified generative models for enhanced image generation.arXiv:2503.01298, 2025
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2d61d451-a3cc-4c58-b64b-c87b07c62c9e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Large language models are better reasoners with self-verification
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 728e9b3f-4e14-4fce-8b4b-57d53df7a461 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 74
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d0cf564-ffd4-4eb8-97aa-2616e011c4d8 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Vila-u: a unified foundation model integrating visual understanding and generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation e81342d4-62d6-4e24-a7e8-b08ab321e5d9 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c0703a0-d95f-4e4d-a78c-9d84baac40c0 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Critic: Learning to Evaluate Multimodal Models
Reference 77
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0121dee9-6d54-4503-89a9-2066e849809e · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 765079fd-8673-4ee6-80de-35b19f652f4c · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61bde09-8bb1-41a4-91b0-14fb9ca232ac · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report
Reference 81
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5292e68c-835e-4792-9d06-564ade5c221a · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling
Reference 82
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 799d2c91-9672-4fc5-8b41-02c3ef15c8cf · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hermesflow: Seamlessly closing the gap in multimodal understanding and generation
Reference 83
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb6049a-a965-430f-a92e-3813bd5b6a78 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4806fd98-19bc-45ea-9e7f-f0b3208da644 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model
Reference 85
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c52337af-fc4b-4ce0-8662-eb2fe8548418 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Language model beats diffusion-tokenizer is key to visual generation
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 643bc847-d865-4585-adc9-b1dbe937335b · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
Reference 87
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97a2c1ed-dfcb-4c1b-9e16-a9853daad063 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 69a51502-6e1c-465e-bf5c-450459cb3dd4 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre-training
Reference 89
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc3c874b-a5c4-43a9-b2fd-aca20f4a9131 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202f015a-ce24-49c9-9654-610a49bdaa9c · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning
Reference 91
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37f6e468-d31e-4e85-a9f6-ac9bc1835124 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 92
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53a839f8-3586-4ffc-af44-54354bb9b2c8 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8fddcf2-8c72-4318-ab50-bdb79916dc93 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 94
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e74cc5c-5dd6-4746-92a7-02b407ac0979 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 95
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4523b3a-873e-4504-abfe-4967ea71a8ad · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning
Reference 96
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 208b340b-771f-4eda-aa3a-b4fde362b2f0 · outbound
UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models
Reference 97
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7529171f-8a64-4c7a-bce7-ccd3c5170588 · inbound
CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e6ede4e-781a-4854-8837-c1962009ed84 · inbound
Show-o2: Improved Native Unified Multimodal Models UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 1fd8eb57-a07d-4e4e-8bcf-eb0a8e181796 · inbound
Reconstruction Alignment Improves Unified Multimodal Models UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Reference 70
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c4b1489-358f-4687-aae1-d500560bbb39 · inbound
Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.