Pith. sign in

Paper Citation Record · LEDGER

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

As of 9 August 2026, this Paper Citation Record lists 73 of 73 outbound references and 26 inbound Pith citation observations for arXiv:2505.23661.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23661 v3

Coverage vector

measured 73 of 73 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:44:18.590800Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 26 of 26 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:37:16.161813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T16:39:58.237442Z

Reference resolution

73 of 73 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved66
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6035d8ec-a76e-4ea3-b87f-954d67ad2cc0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Minigpt-4: Enhancing vision-language understanding with advanced large language models, 2023

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.668718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.668718Z digest=sha256:445b51ee21179d49f43838045fd5d1b07f3945a45dd9f12f1d8d88a5e3b347d0

Observation d4a14748-5626-46a2-b2f0-d7828f34232f · outbound

This paper cites Visual instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Visual instruction tuning, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.783250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.783250Z digest=sha256:0e12c48d35e6646e3e1c1a8427fa99bddb2ab4eed07a050518d077ac8af29860

Observation 578b26f6-e097-41ca-9f47-f8e52869ffde · outbound

This paper cites Improved baselines with visual instruction tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.894781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.894781Z digest=sha256:3b3fe8db40be879e4a1d84e8a966b980009114c1d0092ad9722af64e536bb339

Observation 7054ed70-b2e0-4bd7-8a9a-eef25feca87a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:12.987323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:12.987323Z digest=sha256:b322064179d76739066edbfa6036a9911f5b202869a70fcc5dfcc9b55a1bbc1e

Observation 0d1ff2ac-df0c-4395-b94e-d6748d81253f · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.043657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.043657Z digest=sha256:6cdd81a1f92b5a7ac515ea8ba5564555026eb5079f4220fae6e630e0c0e963e4

Observation 3bf26c07-2194-4332-bacf-f1af516967b0 · outbound

This paper cites Qwen2.5 Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.101067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.101067Z digest=sha256:f7e535d919d757e7b0791e9d3635b6b1f900d9e9d52468ff23a319e09e4f48e9

Observation fd256ac3-fec7-4db9-a8f9-e796a66a9bb7 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.184621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.184621Z digest=sha256:2f5ab80ca2c4640143655225f01c2473f6e52a366ce6de84238a404f9f96f9ce

Observation 48873b2e-ed5e-451d-ab14-b7be49289677 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.275342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.275342Z digest=sha256:613a6b2104fc4ed87e469f45284d186c7c91c21bf9f148414cc8c4db25050fa6

Observation a7932a8d-5df6-4975-9d6e-a08bf667421e · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.368913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.368913Z digest=sha256:f213e555b58e366b52b2878af56fc77b7a4e0d1a6cfe315681e19ddcbae56a69

Observation 3f3ea260-958c-40a7-ba07-2f6de7deef77 · outbound

This paper cites SDXL: Improving latent diffusion models for high-resolution image synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SDXL: Improving latent diffusion models for high-resolution image synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.561728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.561728Z digest=sha256:7b65a8172d087b3933e3371cbd032dad95a7446eefd05bbae023c077573de87a

Observation cc33ce15-5a5a-40cb-91cd-f8ffe9b554a6 · outbound

This paper cites PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.780066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.780066Z digest=sha256:265f23dd1f00b98368bd6898b7076b3e6af83a3bf6e3f6b7dfa3242317bc6f3f

Observation 9ee0b66b-1621-422d-baf0-e6c2443cb82b · outbound

This paper cites Zero-shot text-to-image generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Zero-shot text-to-image generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:13.915135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:13.915135Z digest=sha256:f260f1efb973344f2af831352838654cdd760d095bbc7286d515a6dab0f0f9a1

Observation 4ea54124-5ae1-4e4f-a7df-a3851213fbfe · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.029102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.029102Z digest=sha256:7257e7af1c60e2015b8142b133755301be334ffc3a54caa0f11834032f281091

Observation 8f5966fa-82c8-4c95-b511-30fee3c82857 · outbound

This paper cites Improving image generation with better captions.Computer Science.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Improving image generation with better captions.Computer Science

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.132584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.132584Z digest=sha256:36fdc9bb8b4a33da3f25e9c6ae8a40dfe1d75e19b4071d1743bf4b56c57d342a

Observation 2ed071f9-acf5-477a-852a-ccd165990bc6 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30, 2017.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Attention is all you need.Advances in neural information processing systems, 30, 2017

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.230577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.230577Z digest=sha256:e59a292315d26e72c50d19f9905c46f7cc8f3f4afd6042f8a30166fcff83e101

Observation c87388e8-1a6a-44c0-a5e4-6a5d94ef3bc6 · outbound

This paper cites GPT-4o System Card.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.337586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.337586Z digest=sha256:65edde018cf0041d1340d666bd688e08633438f06cf45f0c35988d792d9077ee

Observation 34ab3c70-925c-4147-91a7-c49f487e21fa · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.494263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.494263Z digest=sha256:5fdff4be44c7bac1b08bde1257397cbac006c6a9afb0ff9145f722e7b2a6ad55

Observation fc3aa865-95bc-4299-a5b2-ae5b1f852c80 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.606841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.606841Z digest=sha256:a4eaff374f2afce54a00824503d858e0c87732c0b8fddb5eab215048bf35163d

Observation a4ff7499-3701-4273-bee7-aaf6c91ab774 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.742428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.742428Z digest=sha256:cc8b43ab89f50e11bdb0fd03d4dd9ba121329466e6916b64531aaf0c6140c56a

Observation feee7fbc-8865-4c9e-a27c-61780c8fae68 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.810243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.810243Z digest=sha256:471a64a28383486741b914a84627708413d7cdf792b1a2da90fce9ecb9c838fc

Observation e20cbcff-517f-4cda-a019-6c68bd4f3b08 · outbound

This paper cites Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Janus-pro: Unified multimodal understanding and generation with data and model scaling, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.891013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.891013Z digest=sha256:1f6e1b6d0a359769e09e1e2d54d4cc6b379f74308b32894fa8d225d874774c7c

Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.957297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.957297Z digest=sha256:ba6367b0c4dabd4570143e3c4a7bc94371b6d5b7be0b622a842335f3ccb794ee

Observation c7590a71-3091-42fb-ac32-7b3bb6302515 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emerging Properties in Unified Multimodal Pretraining

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.039181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.039181Z digest=sha256:3df32ad03a71c0e94ee982ee449b038579248d3a332094050e320ab7b1b70dd6

Observation e9dea184-515f-4a16-8c6f-2a99ab13ca1e · outbound

This paper cites Generative multimodal models are in-context learners.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Generative multimodal models are in-context learners

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.110694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.110694Z digest=sha256:73f5a4b00792e345218f673addee5c0dba97335ddd5a7bf1431894980b82af12

Observation 5c9ce947-b55b-4545-be2e-997b370492f3 · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.148096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.148096Z digest=sha256:8d6e86be07ffcfdd74e00207c588207901d797845b9fc188e4a271a7f5b82223

Observation 189049c2-f882-4d64-b9bb-2f825baa2e5f · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.179360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.179360Z digest=sha256:c2b47fad9a4ceee971de4542a7bf9287af225ecb80bbd8a3ecc26303e18937a6

Observation 10b76f1f-2460-40a6-a315-cdc7ee7d006d · outbound

This paper cites Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip3-o: A family of fully open unified multimodal models-architecture, training and dataset, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.204122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.204122Z digest=sha256:ed180e80bd0bfe75dd87c7845c9a737a419f4bd1eeae0ede9c7ca7f8da961bbe

Observation a5e5cb3f-18c4-41c5-ac11-7305a4af5284 · outbound

This paper cites Harmonizing Visual Representations for Unified Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Harmonizing Visual Representations for Unified Multimodal Understanding and Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.245982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.245982Z digest=sha256:eeb912082937ec2136bac193624b7af452152192b354fea366461fd341197b43

Observation 7fcac129-639a-4db5-b51a-66727b97a7e1 · outbound

This paper cites Transfer between modalities with metaqueries, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Transfer between modalities with metaqueries, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.671731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:15.331127Z digest=sha256:349b627a1ab2b0f92c5f66e6fe115583d1e95eb68c565fdf509e31e9baa275c2

Observation 1c947697-6a78-4fde-b22c-7779b8e343b1 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.407565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.407565Z digest=sha256:d63c5f57016b83f105896c7c201f147d84094eb042b877c04cb98ce4407b0685

Observation 41ef8a11-9e40-44c8-b07d-8ea86291d395 · outbound

This paper cites Learning transferable visual models from natural language supervision, 2021.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision, 2021

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.481465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.481465Z digest=sha256:7d7d12b5ad97eae3bb79068fa1d92a15393d1c5d333e0b4580dbfc2ed225b484

Observation d49b3c90-0a92-413f-a3f3-805b76b068b4 · outbound

This paper cites Sigmoid loss for language image pre- training.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre- training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.558349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.558349Z digest=sha256:95ef74823027b019f35e20e3c6bf5508de2c0732e75bad64f0351839842c418f

Observation eaefd3ea-0f7b-4612-aeaa-352193c3059c · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.611853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.611853Z digest=sha256:321af694b7bba4864d176c8a16b249a994e2f456fd2be511b2f4d83e14443951

Observation 5573576b-e62c-4c13-8f53-a4148241707c · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.658112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.658112Z digest=sha256:861e8b211c9356620ea80614d088f50bcc5b6e181f5485dce53e4b600b92a390

Observation 76af517e-7ec2-40fd-977e-99a915ce80c3 · outbound

This paper cites Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Internvl3: Exploring advanced training and test-time recipes for open-source multimodal models, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.495287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:15.707332Z digest=sha256:b43cea22af5a32845d76c7eecbc385366ede5b07cdcc3d8f4d190f6ac0af889c

Observation 92061989-3c66-4e66-9699-a0e2baa14917 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.737384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.737384Z digest=sha256:124138dd9d027f133aad54dfe380fa8ea43d7e46a3be6eef72d5574ef7b4b6e1

Observation 50d89353-e2df-46d1-9085-6e6aed608311 · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.790154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.790154Z digest=sha256:e4a2d7025fa62c7537d8eb2092a70d785a8a676f47b98b782c8bbd0f34fcdee6

Observation 39b5e017-0143-4b75-9849-cd0125d4f9fd · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.884310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.884310Z digest=sha256:ef25662b2e1ee102429ef094d0c9c625e797a991c3d464330a9c2c97335691c1

Observation 07610958-a4b0-4e79-be63-b8bf11b1e917 · outbound

This paper cites Scalable diffusion models with transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scalable diffusion models with transformers

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:15.937556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:15.937556Z digest=sha256:906ef11b155a34e3333fe2002c8a751231888b56f77050dd2683631d3170fa40

Observation 2077f18a-effd-48eb-be04-092b0ee7208c · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.037523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.037523Z digest=sha256:2be85e59f867d6955171a8ce17eb71273a75904592071f6de1efe07b14506bc2

Observation ec1f466d-522f-403c-9f96-6012750e4d9b · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.103086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.103086Z digest=sha256:2380ee43ef3d4e491923354d85154a24cfdcddf111d9234d4237e2dd4a1012ff

Observation 9ab9c62b-e79b-4b6b-97f6-a13c8b0f88f7 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation U-net: Convolutional networks for biomedical image segmentation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.284865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.284865Z digest=sha256:b1ca86917828f3ce77b321275898c035a33df90918230283eef8909718f4745c

Observation c60d4210-7d90-44e4-b06c-ea52c35bfb5e · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling rectified flow transformers for high-resolution image synthesis, 2024

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.361544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.361544Z digest=sha256:85bbef48e324a6d00021d6ed1d48378f0f0e9f2c09b77a7afad5b2020c9209fd

Observation 5ea1f9f1-6990-4fd9-b2fa-c90bcf99f027 · outbound

This paper cites Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sana: Efficient high-resolution image synthesis with linear diffusion transformers, 2024

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.273660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:16.438696Z digest=sha256:23fa9e1bc2c19b413ec293a8645c29aa1dbc9d47883821afef11e6586acb5cbb

Observation 4dfaf339-61e1-49d1-8b58-eda1097cfc44 · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flux.https://github.com/black-forest-labs/flux, 2024

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.556293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.556293Z digest=sha256:990f6582d6c23e9514423efeac679dbdecc77b3f94211c9fa664ae505168d308

Observation 19350a6c-032d-4ce6-8ec1-3fb0e20ebb8e · outbound

This paper cites Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Sit: Exploring flow and diffusion-based generative models wfith scalable interpolant transformers

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.150952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:16.644623Z digest=sha256:8f2dff8e9df14bf43ed4a119b1eadecacba849b5491f0ed67e2837e11ad98a2c

Observation f46da92d-f9b5-4608-99af-db1e4c6520d2 · outbound

This paper cites Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Lumina-Next: Making Lumina-T2X Stronger and Faster with Next-DiT

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.743577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.743577Z digest=sha256:6cf071d7d7eb51ec82a03f97c3725f4354171c4f7139392e1bc8096bbbeaeeb1

Observation 76282589-2e37-47ca-b250-694a41e03f99 · outbound

This paper cites SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SANA 1.5: Efficient Scaling of Training-Time and Inference-Time Compute in Linear Diffusion Transformer

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.826082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.826082Z digest=sha256:5aa1d50ae7f252c858558abb8310a682243276405da555156b0d837a1f9bd27c

Observation 38a85164-1543-46b7-b883-2d733592397b · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:16.907537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:16.907537Z digest=sha256:fd3dddef2608845b52b36e23a43a9ec0c8f923eff0c6ed96fabcf0ea6ff3c2c4

Observation c8d77a9a-3258-4780-8bc4-d4aadae3a339 · outbound

This paper cites Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Efficientvit: Lightweight multi-scale attention for high-resolution dense prediction

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.012891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.012891Z digest=sha256:a80720cd51e57e7a9fbb27537067e2c88b3e0c5811501e3e55eea9b7046bab4a

Observation e6378ad2-21fc-4c2c-939a-1eaa4ae8a8dd · outbound

This paper cites F-LMM: Grounding Frozen Large Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation F-LMM: Grounding Frozen Large Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.054823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.054823Z digest=sha256:941fe0866d9f756cfc471f7eb0827f86b26c5f53600252643635f0238276a506

Observation 4e7650c6-7a75-4f80-a7ee-22d476024b72 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.119373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.119373Z digest=sha256:18e776f83f5011a9c322e141e4710277a9ce8d72b78e5dca3b0e439c048c808d

Observation 3e059d4e-8c73-46df-a29c-1c8a4bb510f2 · outbound

This paper cites Scaling Laws for Native Multimodal Models.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Scaling Laws for Native Multimodal Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.154880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.154880Z digest=sha256:f0ffc0760712565c44b2be0904c85155b873c3e894c47b302a96e6cf7d92325c

Observation 13c031c1-c6c1-43ea-8e5c-f94f3a17d559 · outbound

This paper cites Qwen2.5-VL Technical Report.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.248177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.248177Z digest=sha256:abab9151390cd844e666935529242de6fa3149a4e5545693408d46a1e79267ed

Observation f667d273-2f2e-47ac-892a-b5ff3481590c · outbound

This paper cites Classifier-Free Diffusion Guidance.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.315211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.315211Z digest=sha256:5b84bcd0b31e03b016cbc1dfd612fddbbda8a72bdd1b481e4b30857185389198

Observation 20ea9665-0bbe-4b92-b2dc-81d09ba090cf · outbound

This paper cites Decoupled Weight Decay Regularization.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Decoupled Weight Decay Regularization

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.387643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.387643Z digest=sha256:c368784a35dc707e59f85021c9996c5919b4c94fd86cbb18fab37d82d9a259f0

Observation bef601de-0ec9-4569-b8aa-0709d22d821b · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.429682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.429682Z digest=sha256:27ea9f266cda5afd2d2cc74731d02501f68f59f425af0bacede529e5e81b6ebc

Observation d648e502-14c5-431f-815f-920dd602f54d · outbound

This paper cites High-resolution image synthesis with latent diffusion models, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation High-resolution image synthesis with latent diffusion models, 2022

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.506293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.506293Z digest=sha256:156b19c24b2c38239a50e2dfe5657fa7368cf70898582dfdf07b6f33bafa1090

Observation 0762f47d-604d-422f-92ef-14c98bdbfb6a · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.567850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.567850Z digest=sha256:4e60787638a643b600695a4b5e5393effc2845fd0f6c474902e381b2495ec4dc

Observation e6d7ba6e-4e7f-484b-8866-57a44e081f07 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.634220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.634220Z digest=sha256:15aaedca6e574f00836e715f8d193040581644e73440e65cff7890b0de942ce0

Observation b128eb78-e9c5-4581-a077-ba2022a1ad64 · outbound

This paper cites Flow-grpo: Training flow matching models via online rl, 2025.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Flow-grpo: Training flow matching models via online rl, 2025

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.685541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.685541Z digest=sha256:4c2bc2f09ea53e7bd59e6456eb721eccb89c7152a693afd983a4a7375f2a9bed

Observation cdcc7b72-102b-4a33-a9df-31ce5984a324 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.764755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.764755Z digest=sha256:1b6140028cf08f9bd9a3dc9c0c714edda861b2cf2a219d2732cfa014e7b65a76

Observation 15e45610-e6f7-481a-aac5-478ccd9f27d4 · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.824218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.824218Z digest=sha256:b16f7de63e4b6b48e05ffa490a36f62fc54a48fa58811fa1999e9c8a8543cb52

Observation 62f640ca-edc3-42c1-af76-465ed47dc62d · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:17.919986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:17.919986Z digest=sha256:3a95b37da0f48350fd84dde1805cefca1e246a63cd6ac03fd46fd9932dcfa935

Observation 010c2984-c358-4ce9-b2c1-2177b04ea6d5 · outbound

This paper cites text-to-image-2M: A high-quality, diverse text–image training dataset.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation text-to-image-2M: A high-quality, diverse text–image training dataset

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:20.017543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:17.986317Z digest=sha256:005b72efb363718fcde10f0566353582fcebc1bb523c4878d7c2f5b3cbcb9430

Observation 538e3db9-6637-4525-a4cb-c950e8931524 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural information processing systems, 35:25278–25294, 2022

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.086468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.086468Z digest=sha256:378ebca96d60dce440973ac3eeecff0544931546d833a118539aeae35ba43f1b

Observation cb5006ae-20b2-4e04-bedc-c6914bae2c8f · outbound

This paper cites Megalith-10M: A dataset of 10 million public-domain photographs.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Megalith-10M: A dataset of 10 million public-domain photographs

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.908808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:18.163452Z digest=sha256:0196cb0ac9012b5310f2b3b498ea84e8d1110cd9cb0e78722f359febf84db04c

Observation c1e9cfc8-45c1-4728-bf5b-8f6682335c2e · outbound

This paper cites RedCaps: Web-curated image–text data created by the people, for the people.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation RedCaps: Web-curated image–text data created by the people, for the people

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:44:19.834909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T12:44:18.256979Z digest=sha256:be7488fd3094f719e5a5acf93a23e5fdc12f273754ddd00b748601b5c8f7d3d7

Observation 44765e00-e05b-477d-8811-bd652baf00a5 · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.323666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.323666Z digest=sha256:1cfb33358e75e39bae7a02246de8dc55e71f52ab54cf0498684317728f11ff91

Observation 0c6db8f2-4aab-47be-90c5-330ac5c78f80 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation Geneval: An object-focused framework for evaluating text-to-image alignment.Advances in Neural Information Processing Systems, 36, 2024

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.427453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.427453Z digest=sha256:b8b72cccec11858c9a02fd5862037a591acac410359cca1f3c65f2c32ee47ba8

Observation eea42e66-daf8-4d68-9b58-bb75b8c65925 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.464184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.464184Z digest=sha256:71a37153659ee756de68c15ab7bf18d6c9e1fadc27ea897f6d200733a573218b

Observation 47103e9b-bdb7-42c9-94ba-44a69aa07525 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.542352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.542352Z digest=sha256:b243faac95d8608b0096fe6438a75f1f03bff636b60b8abcb5f6bfdad934a913

Observation 29dcb26a-57e5-44e0-8b57-e43c8271223a · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:18.590800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:18.590800Z digest=sha256:b8bfeabf18d13e60ac3f98833849a9efd08f4fc79243bd8063c61cf06df4cf6a

Pith citing papers

Observation 6b93d0b8-5e55-41ec-8c49-81e8ff292d4f · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.498867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:7aa1c69ee0ee7adde3899e301db2bfb8bda50ef4284560dbedbd0acb6a1f5731

Observation 2440898c-8e79-4572-8e74-09e812db5517 · inbound

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation cites this paper.

Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:37:16.161813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:37:16.161813Z digest=sha256:1a72ae63e4ce74113429e75bfec29c534de23b684127032a7ef295a25755364a

Observation e6e71b26-80af-437d-8272-7790425ea856 · inbound

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing cites this paper.

Draw-In-Mind: Rebalancing Designer-Painter Roles in Unified Multimodal Models Benefits Image Editing OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T20:06:50.057017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T20:04:00.773443Z digest=sha256:3dbc80525bf0c5a98c2f603839ded9fe7806632ef2aa088d663d2b1b3d0e819a

Observation 2dd93c32-9928-4ffc-a9d7-b37b27e91874 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.243005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.243005Z digest=sha256:a0ba4b6e2ade8f86c2523b2e2ce633731a9b81e838922685b1de863a65de22ca

Observation 1120e7a6-8acc-4195-9c58-944c4448e8fa · inbound

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation cites this paper.

Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T07:03:14.264753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:03:14.264753Z digest=sha256:7c0e6ec7f7cc66864cce5b3225a1341b8a89b14075cad32df24afec283426b59

Observation 21aa7041-6a31-430c-b826-375928f15e2f · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.393690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:3a2da6456fea35149f1ed37fa6cf0ed6fac74c61862def21746a71028e9e871d

Observation c3b208eb-af74-4035-9ed4-2c9adbb461d9 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:52015465d1dd6aeaf2e788893ef4ebe51ec887361358180bf939dd4191bcaa39

Observation 85e93ea3-9f3f-415a-aa55-0158c10054e5 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:21:03.172221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:32:01.551829Z digest=sha256:22d50da3fe7b05d92cd98ac524de41104a9774b4f027ceeaa498b55edb8e46a3

Observation 5b0f89c4-257c-4a73-a01f-98ebefdd4451 · inbound

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training cites this paper.

TorchUMM: A Unified Multimodal Model Codebase for Evaluation, Analysis, and Post-training OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T08:59:55.268707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T08:59:10.877437Z digest=sha256:35a40b8c8699d95565ade74dd16035f6d1f058e2840305f8ff77417d5f00f005

Observation c96aa05d-13c8-4ef0-88af-bcbcf5319edf · inbound

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation cites this paper.

Extending One-Step Image Generation from Class Labels to Text via Discriminative Text Representation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:39.580506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T05:15:22.907880Z digest=sha256:4b632206bdd92be434bfe5264d916c3a68750886a6e49528c297367f11c96967

Observation d8b082f0-af4d-4067-a289-d77dd57f1849 · inbound

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens cites this paper.

Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:24.805355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:58:03.974053Z digest=sha256:67c0a4150f827c0d6dde802ebf49ae7202d5f3a4a5cce6c7e6221f1c1647af17

Observation 519af2ee-a3a0-4291-83e1-c444c03131fb · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:19.050035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:1d3678388ccbe21249a165b68a5b93062792cd11586ec2654f3528b8fb2fd2ef

Observation 921dab5b-a36c-4778-adfb-20bfbf7fbcb0 · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.195508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:0167703255bd2a1ffd517159f168a2640b64ed120ccf3967a7792f2db42a8bcc

Observation ac2fb108-0f9b-431b-8e5c-909ddd734640 · inbound

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality cites this paper.

MUSE: Resolving Manifold Misalignment in Visual Tokenization via Topological Orthogonality OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 146

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:36:07.974566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T15:04:41.518195Z digest=sha256:e7886f7f848eed41b11bfea04e2831506b5623ac7c6fc854dd418015f3d6667e

Observation 333d691b-d2e0-4bd4-9cd5-d6ba23abcb42 · inbound

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision cites this paper.

Steering Visual Generation in Unified Multimodal Models with Understanding Supervision OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:41:10.358702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T14:48:22.805268Z digest=sha256:bdf74f77c3e66bd25778714762a49da0caf8bd81b5a99cffebfd756c3903f091

Observation 008a8304-022a-4fbb-9b37-34e5fbe1fa6f · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.898621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:27067ec51a70887034259a1acedcb29714bf03cba19004b3681cff16592878be

Observation 4473075d-f5ca-4df6-8206-6ce5471a2566 · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.389113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:030413e7c6bad0b41115270f2164f7d2583eedd386de0bbb41684e55b5f6b4ff

Observation 6a177bd3-2801-445e-91e8-250c9704f213 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T11:33:14.301808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:784111051cc53e943c682b41f3368f277ef441f4759730b9fde4614b67ba5d97

Observation a318c7ae-ef0e-480b-9093-0fd3fadac828 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.345303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:7d44c8470b0b194d1ed9e0d7ac4a950f865899edd85655e046c25132b40c13ee

Observation 56505dc3-693d-40eb-bc29-692615926173 · inbound

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers cites this paper.

HYDRA-X: Native Unified Multimodal Models with Holistic Visual Tokenizers OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-03T14:38:28.896947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T07:01:07.362430Z digest=sha256:f33eaa2de3e490fdf6e637746c2d973a9e9a78f37d7a3a3f5117b784c4acf42b

Observation 8243c92a-4a9d-47b3-a384-fee137bb55cc · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.697635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:08:25.661515Z digest=sha256:172bc54dad727ededac7ab3fe99d9b9ae6596ba6e549ad4bc704cb50303e3d73

Observation 2e827ec3-e9be-4230-b11a-982702ef9b13 · inbound

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models cites this paper.

SPAR: Semantic-Pixel Self-Alignment and Adaptive Routing for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.439470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T23:15:09.253879Z digest=sha256:cc0c99b5f0224d5f863f8cd74652db3959507b8a029886f4a9f98ac96c645282

Observation 962dcbc7-89cc-4983-8fef-605c13eafe22 · inbound

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation cites this paper.

IV-CoT: Implicit Visual Chain-of-Thought for Structure-Aware Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:39:58.238735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T00:19:49.071495Z digest=sha256:8deb3689ad4ecdb5867d56101117e1e023e57727d86cf5ace9e2e69bfdfe9d4c

Observation 4335c52f-26bc-460c-b559-203232a2b955 · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-13T05:11:19.001184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:11:19.001184Z digest=sha256:7d8c41fced76d8f8046611b7c200ffafcc85bf78835043a0fafbaff12bcefa68

Observation 8ae6621d-7e99-4606-8967-b66ff2bf033f · inbound

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation cites this paper.

IB-Flow: Information Bottleneck-Guided CFG Distillation for Few-Step Text-to-Image Generation OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:18.617071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:18.617071Z digest=sha256:8e012fa941c8a63e145d67d7e6869fe93ffedd9dae1fac50d448cb6d1421e2e6

Observation 5d74a6e2-0343-4150-a5f2-213031a8a0d0 · inbound

Test-Time Curriculum for Open-Set AIGC Detection cites this paper.

Test-Time Curriculum for Open-Set AIGC Detection OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-05T00:46:23.340701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:46:23.340701Z digest=sha256:7a5b32437e6116b729db80130dba148172c5b3ed8fb1a6a2145994c84ebcfe5c