Pith. sign in

Paper Citation Record · LEDGER

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

As of 16 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 27 inbound Pith citation observations for arXiv:2505.23661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23661 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:18.590800Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 27 of 27 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:06:24.629360Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.237442Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6035d8ec-a76e-4ea3-b87f-954d67ad2cc0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.668718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.668718Z digest=sha256:339e6cfd857cef001017156e6092e03cac4ff174b3969427073b7ac68108c99e

Observation d4a14748-5626-46a2-b2f0-d7828f34232f · outbound

This paper cites Visual instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Visual instruction tuning, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.783250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.783250Z digest=sha256:36a98dd5b920302d6783226523db65b1da5743bb7c3526466b84fbda740ff359

Observation 578b26f6-e097-41ca-9f47-f8e52869ffde · outbound

This paper cites Improved baselines with visual instruction tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.894781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.894781Z digest=sha256:add7830cff3e5820c85285c52c59f04f731880951762b108cbc7fbc0ce84b7cc

Observation 7054ed70-b2e0-4bd7-8a9a-eef25feca87a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.987323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.987323Z digest=sha256:553efd4e854f297755d6975c24128cf062ee39d66382a450ea767ab564530ce8

Observation 0d1ff2ac-df0c-4395-b94e-d6748d81253f · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.043657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.043657Z digest=sha256:4649c3f13bdebc9581de104427c599cffa61cfb0b4eb18d611c616afa0962d62

Observation 3bf26c07-2194-4332-bacf-f1af516967b0 · outbound

This paper cites Qwen2.5 Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.101067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.101067Z digest=sha256:d5e26e8e60c72a99b21754cad17cbe5806da260823e5cdb0f2457b7bb198c980

Observation fd256ac3-fec7-4db9-a8f9-e796a66a9bb7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.184621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.184621Z digest=sha256:c14d2cf47ff403415ee859409027670b24af0bf4bb944c3aab2c79ab60777a4d

Observation 48873b2e-ed5e-451d-ab14-b7be49289677 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.275342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.275342Z digest=sha256:6bbc074a044d3031c48ac6939fab2dd2f8399b2a5100f41bf1978d1b353db6f6

Observation a7932a8d-5df6-4975-9d6e-a08bf667421e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.368913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.368913Z digest=sha256:60060a4cd560c362adebb6483458f97e737e13bbaf19271768c025fa0dbd4d67

Observation 3f3ea260-958c-40a7-ba07-2f6de7deef77 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.561728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.561728Z digest=sha256:3355f361112ef3a53ee2c6d0706866d08eecf20f82fb1ba7e7a5ffe7b6e89a81

Observation cc33ce15-5a5a-40cb-91cd-f8ffe9b554a6 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.780066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.780066Z digest=sha256:336374bbf19316e81a065bfe68f63028ad54bd7329831483cc803863e7ea17d6

Observation 9ee0b66b-1621-422d-baf0-e6c2443cb82b · outbound

This paper cites Zero-shot text-to-image generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.915135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.915135Z digest=sha256:3c382636233d4497a6f1fdff8ce2e394d1a13c08c59d1eeae3f668c2ea50985f

Observation 4ea54124-5ae1-4e4f-a7df-a3851213fbfe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.029102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.029102Z digest=sha256:915e61d8fee0347df6e8c30eaccaed15dd6cbd6b0f8bf613f0fcac579ec4e69f

Observation 8f5966fa-82c8-4c95-b511-30fee3c82857 · outbound

This paper cites Improving image generation with better captions.Computer Science.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improving image generation with better captions.Computer Science

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.132584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.132584Z digest=sha256:d94e8934a2ab94c9224f692ce34a1f475ab399b2b5cfbc30bc43708bff6b0f2e

Observation 2ed071f9-acf5-477a-852a-ccd165990bc6 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.230577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.230577Z digest=sha256:079dff5e9d120bcc53dd3badd8f4d6a6b5a5dfa604665c82b650a9f3c4ef7d5d

Observation c87388e8-1a6a-44c0-a5e4-6a5d94ef3bc6 · outbound

This paper cites GPT-4o System Card.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.337586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.337586Z digest=sha256:0ab95d41933c40b5897f05fd789b9b0cd411c4882eb48f8ed85a1bdf83f85ddd

Observation 34ab3c70-925c-4147-91a7-c49f487e21fa · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.494263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.494263Z digest=sha256:7914b151df95f10091531a802c10c566ac06450e909585be52818be59ee752c5

Observation fc3aa865-95bc-4299-a5b2-ae5b1f852c80 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.606841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.606841Z digest=sha256:b61dcb9337ff9a05001c304de8607b1bd6e94731e90f036f6cbfd91dc1105f80

Observation a4ff7499-3701-4273-bee7-aaf6c91ab774 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.742428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.742428Z digest=sha256:bedeff468504640a576b3c698fc61dd49ae8f40a5de57368c46e775bad391f31

Observation feee7fbc-8865-4c9e-a27c-61780c8fae68 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.810243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.810243Z digest=sha256:cbb645bc7ccd7d889105856cf216328a31c2e4e8673b22fd4d48d6d07b560998

Observation e20cbcff-517f-4cda-a019-6c68bd4f3b08 · outbound

This paper cites Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.891013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.891013Z digest=sha256:c11ca0170a0865566b57e52964beac5b430ba0767170553c92ef85ade0049817

Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.957297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.957297Z digest=sha256:0ba7df32ac1eafda35b541fb13bc74490268c09295dee872811eb8e22687c916

Observation c7590a71-3091-42fb-ac32-7b3bb6302515 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.039181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.039181Z digest=sha256:944d88d41e9e9d6e56a85bc56fc2aee2867c11e7358a97627c3584fc7c4a0f82

Observation e9dea184-515f-4a16-8c6f-2a99ab13ca1e · outbound

This paper cites Generative multimodal models are in-context learners.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.110694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.110694Z digest=sha256:27bd979399a5022f176d9e19d3184bef8be6ba60c042289e21a062e965c6cb0f

Observation 5c9ce947-b55b-4545-be2e-997b370492f3 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.148096Z digest=sha256:c1a5abe5166f3340f2684411057802bb75dec301d89516ae256b38fe5ec5c1cf

Observation 189049c2-f882-4d64-b9bb-2f825baa2e5f · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.179360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.179360Z digest=sha256:0af80697c68dc54dde1a58da1090d84f977af7595c3a091c5bbaa0280e1321bd

Observation 10b76f1f-2460-40a6-a315-cdc7ee7d006d · outbound

This paper cites Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.204122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.204122Z digest=sha256:189d1e301fda29c5ab9315b491273a1e90964826e9c90d8d47c7c25a4622c3be

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:ad95cd92fbd75be7b27d96bfb902b98339cbb93c8cef264903314341454609ee

Observation 7fcac129-639a-4db5-b51a-66727b97a7e1 · outbound

This paper cites Transfer between modalities with metaqueries, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Transfer between modalities with metaqueries, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.671731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:15.331127Z digest=sha256:ce1cedbc8e70d4c5e6d76f929b3c88a1fa287ebff3ee0adede87fd77457b4f6b

Observation 1c947697-6a78-4fde-b22c-7779b8e343b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.407565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.407565Z digest=sha256:049d82a7e6c0d7ccee2042f21977f34cdf3d24596f270c12b91cdcd48c380f12

Observation 41ef8a11-9e40-44c8-b07d-8ea86291d395 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision, 2021

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.481465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.481465Z digest=sha256:2f0d9ffba13b3525381c6bd85af4122bee82d6bfd5f8edef0671a1f019228189

Observation d49b3c90-0a92-413f-a3f3-805b76b068b4 · outbound

This paper cites Sigmoid loss for language image pre- training.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre- training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.558349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.558349Z digest=sha256:5d81870a14b5405df74a867f5b4ba61d2c0e483978cfe943f6c7a472cf966b4a

Observation eaefd3ea-0f7b-4612-aeaa-352193c3059c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.611853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.611853Z digest=sha256:e3795fd0ecf9484397eb6f3ebf241f804f11e2908662239aa2b51b6f5f58e652

Observation 5573576b-e62c-4c13-8f53-a4148241707c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.658112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.658112Z digest=sha256:7f2a5cc2aa698d9ed207c3af01aee5251bc35e1cfd76339e1de57403e02946c3

Observation 76af517e-7ec2-40fd-977e-99a915ce80c3 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.495287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:15.707332Z digest=sha256:e19546146fc89a68e8372c9f2611ab695f1c3c03f5fcb3cc0e9619770a439d51

Observation 92061989-3c66-4e66-9699-a0e2baa14917 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.737384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.737384Z digest=sha256:87c67bab48cf76a477ce003183e2edb82898bcd6e3b3117e6f036412059f54b6

Observation 50d89353-e2df-46d1-9085-6e6aed608311 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.790154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.790154Z digest=sha256:5023ffe216314bd46f92e48bf76935576fcccc8b9dd62475c3f0b98312f2d2db

Observation 39b5e017-0143-4b75-9849-cd0125d4f9fd · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.884310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.884310Z digest=sha256:ba8a4a732d8b83863811de56cbf2e4f77023809510501bd0ede0ad05edd0f6ca

Observation 07610958-a4b0-4e79-be63-b8bf11b1e917 · outbound

This paper cites Scalable diffusion models with transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.937556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.937556Z digest=sha256:7f1a71b50d09ca86a9c591f0a6f4aac6607131dbd258a2e0a6bd58d21b8cb793

Observation 2077f18a-effd-48eb-be04-092b0ee7208c · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.037523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.037523Z digest=sha256:d9d7c3a2183e270f9363ce986f4482823230920b469c61fc41c601f26cbd3611

Observation ec1f466d-522f-403c-9f96-6012750e4d9b · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.103086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.103086Z digest=sha256:8d55c0ba554077c08e287a35a99a1147b2e4e556cfd45dd24aa97443fde197f7

Observation 9ab9c62b-e79b-4b6b-97f6-a13c8b0f88f7 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation U-net: Convolutional networks for biomedical image segmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.284865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.284865Z digest=sha256:1d27cf35f5e2994970a61abad19458158f802c71e8c5caaf3494036017393e8c

Observation c60d4210-7d90-44e4-b06c-ea52c35bfb5e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.361544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.361544Z digest=sha256:45fd2b0970b9a220411955da43f78644f2ee39fea06a331ab79efce9f481da5b

Observation 5ea1f9f1-6990-4fd9-b2fa-c90bcf99f027 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.273660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:16.438696Z digest=sha256:ec16dce644984a9618bf23fc548b5b1dda9a23e84ec25877470467b655678c4e

Observation 4dfaf339-61e1-49d1-8b58-eda1097cfc44 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.556293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.556293Z digest=sha256:436cdc2c42668ad6bbaef1b67d89749e15f0379207775399c0125fdfc855464c

Observation 19350a6c-032d-4ce6-8ec1-3fb0e20ebb8e · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.150952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:16.644623Z digest=sha256:a11fedc92b99761fc355ffebad38a444bd4e148b9ec7660f65335746bb83b48d

Observation f46da92d-f9b5-4608-99af-db1e4c6520d2 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.743577Z digest=sha256:d41d9d1a0af3b7b86463b293f4391d3c626b883ce0d83f3b97a72be3b707419b

Observation 76282589-2e37-47ca-b250-694a41e03f99 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.826082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.826082Z digest=sha256:7c4128a9f78f6cad743a7c3b060612ebb400ac440eff37998e4c116c18591e8f

Observation 38a85164-1543-46b7-b883-2d733592397b · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.907537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.907537Z digest=sha256:b84bc1dda447fb57af34744597be205efd51df0e43a74268400a8097b580b1ac

Observation c8d77a9a-3258-4780-8bc4-d4aadae3a339 · outbound

This paper cites Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.012891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.012891Z digest=sha256:7516d1dceafbc3648fb83e5e57aed8dc4a215b4ed91d648c1654bd16ca934e29

Observation e6378ad2-21fc-4c2c-939a-1eaa4ae8a8dd · outbound

This paper cites F-LMM: Grounding Frozen Large Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation F-LMM: Grounding Frozen Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.054823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.054823Z digest=sha256:edbf9e59c5bc51b496d319cbd8c364bc89f0f7f41a813b12f9f5700445f20010

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:4b8dad2ac40fc3b6274a221d6ca63c5b36a2e777a29d191cb641fe72445a06a9

Observation 3e059d4e-8c73-46df-a29c-1c8a4bb510f2 · outbound

This paper cites Scaling Laws for Native Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling Laws for Native Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.154880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.154880Z digest=sha256:38ac8bdcccd3401a7fd6b29af3e1065220a38cbbd9c26ed33cbd2cf1857badb7

Observation 13c031c1-c6c1-43ea-8e5c-f94f3a17d559 · outbound

This paper cites Qwen2.5-VL Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.248177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.248177Z digest=sha256:8a4e1d2b964cb745e4dc056f4b1d1c9e0759a79c94a22bf404dbf8ed06caa302

Observation f667d273-2f2e-47ac-892a-b5ff3481590c · outbound

This paper cites Classifier-Free Diffusion Guidance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.315211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.315211Z digest=sha256:9c098e473149bd011cb83f36bb13aae3be706cf7eab031fe2a3c34b540b672f7

Observation 20ea9665-0bbe-4b92-b2dc-81d09ba090cf · outbound

This paper cites Decoupled Weight Decay Regularization.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.387643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.387643Z digest=sha256:9e1824aebd17b63cb25eca56d1eb1eadd9cf59bce2464e34a73f3ecadb95276d

Observation bef601de-0ec9-4569-b8aa-0709d22d821b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.429682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.429682Z digest=sha256:534b5eb7d9b4582afbfacbfa0b31ccfc305297b8ad349d93e74a622597bc2b8f

Observation d648e502-14c5-431f-815f-920dd602f54d · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models, 2022

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.506293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.506293Z digest=sha256:e8121680826ccb94f53863ca9f32f52c3553293fd639040f395cd50ef2b7a7c9

Observation 0762f47d-604d-422f-92ef-14c98bdbfb6a · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.567850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.567850Z digest=sha256:89568dabaceef810e139558f8558269436878fa36af729c58742978e6674e56a

Observation e6d7ba6e-4e7f-484b-8866-57a44e081f07 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.634220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.634220Z digest=sha256:1d16f7014c0704016cfeeda8cc97ee86f4a3bd69e66ac6bd7e3d31030294f08e

Observation b128eb78-e9c5-4581-a077-ba2022a1ad64 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flow-grpo: Training flow matching models via online rl, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.685541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.685541Z digest=sha256:4d11824d77e9cb9769bd553ce4db57fc4fd3cc1a87ce23184977acd5c02aa7c6

Observation cdcc7b72-102b-4a33-a9df-31ce5984a324 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.764755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.764755Z digest=sha256:0fd856760fe928ad456fc1e1ebf59c5f20e76453bf7f5dc9f2ccc85fef792501

Observation 15e45610-e6f7-481a-aac5-478ccd9f27d4 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.824218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.824218Z digest=sha256:9babcd75cd7a6f308bb30d828e12b3ecffa99bfd06874ddc8c3834c25d7068f1

Observation 62f640ca-edc3-42c1-af76-465ed47dc62d · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.919986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.919986Z digest=sha256:4d52015759c233a727ac8370f40e6a0bf5e7a65bd6ee2ff0c201512129b924a4

Observation 010c2984-c358-4ce9-b2c1-2177b04ea6d5 · outbound

This paper cites text-to-image-2M: A high-quality, diverse text–image training dataset.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation text-to-image-2M: A high-quality, diverse text–image training dataset

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.017543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:17.986317Z digest=sha256:1c8b5cabe1edda1a358836e50d64d71023efd0767203b89b890f0ab42f6883d4

Observation 538e3db9-6637-4525-a4cb-c950e8931524 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.086468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.086468Z digest=sha256:113f40882c3420d9ce59159efaa9048fbf9801d4ae97f38daf4dffc213edfa12

Observation cb5006ae-20b2-4e04-bedc-c6914bae2c8f · outbound

This paper cites Megalith-10M: A dataset of 10 million public-domain photographs.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Megalith-10M: A dataset of 10 million public-domain photographs

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.908808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:18.163452Z digest=sha256:6d036ea36f41c097b26a53950d10a7e8b0bbb6cf9f5d8fb1863c46b37f0a35e9

Observation c1e9cfc8-45c1-4728-bf5b-8f6682335c2e · outbound

This paper cites RedCaps: Web-curated image–text data created by the people, for the people.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation RedCaps: Web-curated image–text data created by the people, for the people

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.834909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T12:44:18.256979Z digest=sha256:c5d9751d0fca0a92badc8cb1f1e74526da90343c25bf9a0bdc129b63e121a171

Observation 44765e00-e05b-477d-8811-bd652baf00a5 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.323666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.323666Z digest=sha256:250227b1ff7c6a9897d469e3a6db5b96bd6c01b878e64c7211a2d45fea8cbd62

Observation 0c6db8f2-4aab-47be-90c5-330ac5c78f80 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.427453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.427453Z digest=sha256:a75da6add98852a94e95ea5db996e3cf73ed7484f5faf86e7c55391094fabaec

Observation eea42e66-daf8-4d68-9b58-bb75b8c65925 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.464184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.464184Z digest=sha256:59dee024821cf6fdf471c87394af417d2854ded11c60ca0ea61c7e22838b99ad

Observation 47103e9b-bdb7-42c9-94ba-44a69aa07525 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.542352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.542352Z digest=sha256:235cd009facabce0a1511b9882fb57925c5dfbb98b9ba37155ef7224c3b1b174

Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.590800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.590800Z digest=sha256:c8bfd671c43ef0d3ca2b807e15f269a9ca2ee869e6ca4f39e301805e8528fb22

Pith citing papers

Observation 6b93d0b8-5e55-41ec-8c49-81e8ff292d4f · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.498867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:426d6de3c361a1a67acd360f5e5c5a83149e2358df5c47ca38e313004bbd7974

Observation 4dd38f45-424e-4877-86b5-f8c11c78c482 · inbound

SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer cites this paper.

SA-LUT: Spatial Adaptive 4D Look-Up Table for Photorealistic Style Transfer OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:06:24.629360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:06:24.629360Z digest=sha256:c1056d2bc34801a479bd84e5d616cad9d98bbd6356f073934ef3a56971c3e9b9

Observation 2440898c-8e79-4572-8e74-09e812db5517 · inbound

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation cites this paper.

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:37:16.161813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:37:16.161813Z digest=sha256:ab7f4bdafafb5dac6134cf2db7bcef1ebce6ba80ae02c8fd6496b23f3550a558

Observation e6e71b26-80af-437d-8272-7790425ea856 · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:50.057017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:9f08164f66c2e03b7fa5515cec2a84c4ffc13d47fd9c5e122066de345249abcf

Observation 2dd93c32-9928-4ffc-a9d7-b37b27e91874 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.243005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.243005Z digest=sha256:1beb9ea30aed65408702ac5fe93c61696565a44eb743711d16e96bb249cba438

Observation 1120e7a6-8acc-4195-9c58-944c4448e8fa · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:14.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:14.264753Z digest=sha256:0a42910d92dac834078006a04a83c93060e3a5a3fc8d9aaa908d30c1380d8997

Observation 21aa7041-6a31-430c-b826-375928f15e2f · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.393690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:2eb23c074c6e1d0181259931f7caa8cf6832b6af85b7b94806fee1959b429577

Observation c3b208eb-af74-4035-9ed4-2c9adbb461d9 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:3d49a1c5c93d9fa8f71b71386762b99c001b5de348850e74b027986f3afd042b

Observation 85e93ea3-9f3f-415a-aa55-0158c10054e5 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:03.172221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:43f4bb4985a138fac112f00953f0c2b5733e8fbbf59ab922c0e7cb8f43b6431f

Observation 5b0f89c4-257c-4a73-a01f-98ebefdd4451 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.268707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:a5f23d408fce46fa569068e1db09d2154b813990be4cf499d994765ec34b770e

Observation c96aa05d-13c8-4ef0-88af-bcbcf5319edf · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.580506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:bb699d891fe90cc8c2d652cb8853df2e5421a9dfa372b573604149cb067f028d

Observation d8b082f0-af4d-4067-a289-d77dd57f1849 · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.805355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:c8b7a1006f5f02da6f11f7a21de19ae65dffb56d9e3992ff211b844959fccdf6

Observation 519af2ee-a3a0-4291-83e1-c444c03131fb · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.050035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:5d6e1c2bc2c6ae5550a14a3f314e16dcfe27e6e86edc36501cc452180d370a37

Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.195508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:2ebe4dbc764eb4c013f8310c84990566d3b74d8cd77702451f723c9bd0378bf5

Observation ac2fb108-0f9b-431b-8e5c-909ddd734640 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:07.974566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:90d262bc24b354cbd4b6abeb7665d56872d4acd9870cdccb35838f6165554e24

Observation 333d691b-d2e0-4bd4-9cd5-d6ba23abcb42 · inbound

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision cites this paper.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.358702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:9e19ef8711d38ff5d4bdeb083eb394dcbbea092a5ba52761dbbbecbf24decba8

Observation 008a8304-022a-4fbb-9b37-34e5fbe1fa6f · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.898621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:babf30c387c67eb3496100bfe9efca6ea069c7a8fce42b416601a75b8acb61dd

Observation 4473075d-f5ca-4df6-8206-6ce5471a2566 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.389113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:bec4d795f6473a0ca1bafbb73cd334684def1749e96313ae655f10e84c0b9680

Observation 6a177bd3-2801-445e-91e8-250c9704f213 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.301808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:7e64297efcaf36c4caa975dec72b5bb923588ef7b00dd525b3a807fc6e3973cc

Observation a318c7ae-ef0e-480b-9093-0fd3fadac828 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.345303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f9116d878438e004aca350794925956e3c793ea4b37e34640f2745ddd3e75586

Observation 56505dc3-693d-40eb-bc29-692615926173 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.896947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:7d4bec9fb278704d5181867a78828d5affee3275ae53eeab289e5a0c03079197

Observation 8243c92a-4a9d-47b3-a384-fee137bb55cc · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.697635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:107159a211c1d148a235971b1b4b28da6c148b64207c9e73095c84615de5ed6a

Observation 2e827ec3-e9be-4230-b11a-982702ef9b13 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.439470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:6b1fbbf4979e6db4cb03091e32150f188b52c9f826feb8e01c81479603d74334

Observation 962dcbc7-89cc-4983-8fef-605c13eafe22 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.238735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:bbcc036f038821391304d071ca878911d63a1ab7cde0ec1e895e9877531d5d06

Observation 4335c52f-26bc-460c-b559-203232a2b955 · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T05:11:19.001184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:11:19.001184Z digest=sha256:dbc46f0e519960c5a30854ae92c65393d14ccc11faf0e67ff6130463b97f7eab

Observation 8ae6621d-7e99-4606-8967-b66ff2bf033f · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:18.617071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:18.617071Z digest=sha256:46e41654dbeb2eac3551339a74f27b474e4be193304ab8f7e2b98423e9a27443

Observation 5d74a6e2-0343-4150-a5f2-213031a8a0d0 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.340701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.340701Z digest=sha256:e164731eca2fce62938c9a2c7fe70649805136a15be642d065da381560497db8