Pith. sign in

Paper Citation Record · LEDGER

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

As of 7 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 7 inbound Pith citation observations for arXiv:2505.23380.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23380 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:53:27.275816Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:36:08.089323Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T13:39:51.505020Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy6
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aec5fbc9-7aeb-4bbe-a4ec-b481b7096a02 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.247048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.247048Z digest=sha256:561968f1f053c5bb3970e88741c8af67abbfa3af6f78d0e0e15251cd09ffde5f

Observation 58bb2215-a757-4c2a-965b-9064c404f0fa · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Gemini: A Family of Highly Capable Multimodal Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.392275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.392275Z digest=sha256:7751f097a94cbcdfe0cda0be64ec5d505751f43158cec8e7413be4b52eb75d18

Observation b548974f-40f1-47cd-8543-216fa5d1fe4e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.498548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.498548Z digest=sha256:f3c418fc0a5fe58dcdafe9900cfb034852e6cbd02c478081e0e1ecc11f894389

Observation de27a906-4d36-4c4a-a3db-7a66dbc3c57d · outbound

This paper cites SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SFT or RL? An Early Investigation into Training R1-Like Reasoning Large Vision-Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.647902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.647902Z digest=sha256:8fd0207e2bc03ba6eb5d986387e5e91a09ac40c998d53386b9a9c6d8d35d9832

Observation 5188efac-6b8d-4830-88d7-73269d8607d5 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.803824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.803824Z digest=sha256:79f45d123364790641737ca8711ae39e1c8ccdd290c4d9772f7057d9a68cbb51

Observation 451f3035-a90e-470e-9aec-921385dd5ba2 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:21.903111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:21.903111Z digest=sha256:17d9cb828bf10bc6a20d95419a49e455086b9a23e03b048f0e20fa93e397186e

Observation 20b0f5b4-8f58-44e8-a842-3f91de0cca06 · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.932772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:21.990830Z digest=sha256:f8abfdf7d00a578d61ff5090ab11646abb249afb3b94b6af36701fc1f330ba49

Observation dd42ec44-cf96-4338-868e-c364240aa56b · outbound

This paper cites Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Baldridge, Roopal Garg, Peter Anderson, Ranjay Krishna, Mohit Bansal, Jordi Pont-Tuset, and Su Wang

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.653559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:22.106661Z digest=sha256:ecfbd332942d9ad56f77ed075a68b7613831ca0b4df228b8fff8ab8f4755641a

Observation 9dc2eece-21b3-4a7d-a269-9cd6afc167cb · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.203748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.203748Z digest=sha256:c0bc765b86d77c0e0650a5e78524a06e18f6d56e6abd37795c9d161f70d661dd

Observation b9db14ca-f33c-4c9c-a4ac-1bcc3f660906 · outbound

This paper cites DreamLLM: Synergistic multimodal comprehension and creation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DreamLLM: Synergistic multimodal comprehension and creation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.295884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.295884Z digest=sha256:8437f05f7416ac1ae3d9c29217f1b8f5c1a57627e5a198f311a89bf4c3cbbf68

Observation c8db548b-d4fe-41f3-a841-c5bd44d5cf05 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.410352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:22.391653Z digest=sha256:38ef0e03084469fb37ccb3e139c566ba630938f1575e7e5db8071c17709ea777

Observation ba29a64c-395e-4fe1-afd7-b2d004a890d3 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Taming transformers for high-resolution image synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:30.134836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:22.481332Z digest=sha256:b9e2a0a4ac83472e74831f39eda9efd6365c090361c8f4ad68a7820163934bee

Observation 66ebf948-5046-4075-86b9-878955536ed3 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.622402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.622402Z digest=sha256:74e7c8296544e3a52380020cc7b6c226c240dad2bd5fa20073c8996bf9ea965b

Observation 018ec1df-5998-4250-a5a9-2177b76fca32 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.801107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.801107Z digest=sha256:aa05d1771a5612162956456f1e6dbaaa4bd3232c0797de786a3ddceed26ac538

Observation 3e8053ea-4c2d-445e-aebf-e7b2b1ee8a09 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:22.895103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:22.895103Z digest=sha256:e3212b5e829cb4b8a6952760a2410f52c54dbc06919fdcfd451b0342e8d711ee

Observation 9b0265b6-958f-4093-95cd-96faa8adc758 · outbound

This paper cites OpenAI o1 System Card.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.010341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.010341Z digest=sha256:0a64574eb01149274258a077e8605f73220e489557f4fd895a143c2f2633a0e3

Observation 668be6c0-d07f-4c16-b854-b3f4c42bb6ef · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.148041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.148041Z digest=sha256:914da8746c38be21eeb141c7134ce37642dcb89ab52373fddab3f8b22cae2eb3

Observation 2c634cf9-a721-40d5-8aac-613f2429c5f9 · outbound

This paper cites Imagine while Reasoning in Space: Multimodal Visualization-of-Thought.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Imagine while Reasoning in Space: Multimodal Visualization-of-Thought

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.280824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.280824Z digest=sha256:d3d600e580d6e1b8df26928a64a57782b1203e19cbbc5a68fb21232716ad143a

Observation d4fa1d2f-af8a-4a60-aea4-08ba69542773 · outbound

This paper cites SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.378411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.378411Z digest=sha256:fce0bf42d7731dbbfa56d4dda9b1af5b5226beecb640979c38ca846c90ef499b

Observation 43be3b75-d5fa-411d-a840-991ebc006b4b · outbound

This paper cites OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning OmniFlow: Any-to-Any Generation with Multi-Modal Rectified Flows

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.481480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.481480Z digest=sha256:fb9f07b3e0aa8d993f4780c889d3430c9deef99953f76ad77290fddcee03ea05

Observation 2e9ad2e1-f710-4333-9696-fce27c100765 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Evaluating object hallucination in large vision-language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:29.928449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:23.546372Z digest=sha256:82c5dc54a5c278f42ed185b8c37d8de0966eca4cd54a32b42de1571f3fe8f238

Observation fc666632-e083-4888-866b-1ae50628efdb · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Textbooks Are All You Need II: phi-1.5 technical report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.636560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.636560Z digest=sha256:2dbdee1c87835bf419bb58fdc2a21cd026da4d7147b6ad2cb1a31cd13bc90459

Observation c72fe73b-cab8-499d-8243-8f62cb7fce89 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.757213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.757213Z digest=sha256:2122bcf18ee3383edbeb98abd1a74afcf395708246c33816156de34db15e0068

Observation 51de8089-3bda-4734-b1ce-cfd61af89862 · outbound

This paper cites World model on million-length video and language with ringattention.arXiv preprint, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning World model on million-length video and language with ringattention.arXiv preprint, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.884539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.884539Z digest=sha256:7420856123305279f5da3d6b2479dbe5552d115ee90def99c1f60f56f9c7674e

Observation ef2e40da-f413-46e3-91e8-cafa6f0483eb · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.079598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.079598Z digest=sha256:94652342433f417e20e991b44934f0aab835b35c2b58b7bfb50cd946f49a1f8b

Observation cd8acce9-ac40-4322-a4be-1e017b220653 · outbound

This paper cites Visual instruction tuning.NeurIPS, 36, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Visual instruction tuning.NeurIPS, 36, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.253921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.253921Z digest=sha256:271f76711175c4839282bed27c842b4482bf1cd48ef6d88b6009569af51eef71

Observation f7a0385d-c4bf-4561-aeb5-f7d4665e6dd0 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.426795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.426795Z digest=sha256:a9c9496aafcf31e71131706bb5a8e8d75f6de3385670fc7ea79a23dbd7c1110b

Observation abbcd6f8-fb9e-4226-bd62-53b52b9d3f5c · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.588515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.588515Z digest=sha256:12f015387f2c5cb21978cc538d44c67086a039c2459e82b8b59ee71993dbfcbc

Observation 8ed1468c-7cf9-4a0c-b1c7-8dcfafffb433 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:24.689189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:24.689189Z digest=sha256:3a30b3a8aa2771611489e8175c5ce004f72546a53346ed04f5c7e00f1a7aa829

Observation f7ab2202-5522-41f1-b062-d5223aef3e5a · outbound

This paper cites Learning transferable visual models from natural language supervision.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Learning transferable visual models from natural language supervision

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:53:29.228511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:24.823422Z digest=sha256:13b57c68bc645312c7b55e075ad616d1c82ef5924bc02f34585791941a91ec82

Observation dc850bae-8438-49fc-b1eb-d0faf2e5271d · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Manning, Stefano Ermon, and Chelsea Finn

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.018029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.018029Z digest=sha256:e7e010bc53cc8ec4ddebc055cdc88363584dbcee24ebf0fc59b9eb027f8a8966

Observation 971cb1b3-6d1f-49e9-bc2e-e50d243c1c7d · outbound

This paper cites Proximal Policy Optimization Algorithms.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.202348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.202348Z digest=sha256:30d75bb5ce67152a385bb83982b1a8b3d728c1b00b067e437037080d30b1952b

Observation 62ac2f3a-ce0f-4374-8b79-36b0b4b1a2de · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.336876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.336876Z digest=sha256:65c4b4e8e6487db01e35625bab33462384374da3de34fef2567d0dec847cbc6f

Observation 05dac12b-2ec1-4114-b49e-baa9bdda25cf · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.463377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.463377Z digest=sha256:e6bb2193864e45445949c7b1fac542b430693d0b365da190b13bda620240c817

Observation 2b70c008-595a-485d-b86a-263d5d806029 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Journeydb: A benchmark for generative image understanding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.595623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.595623Z digest=sha256:252dc7544b798f9314bc595e953b7e35f1327cbd737563e0659834a070fb3386

Observation 16b7d5a9-ffd4-40e9-a283-5569a2b83d4b · outbound

This paper cites Emu: Generative pretraining in multimodality.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Emu: Generative pretraining in multimodality

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.723242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.723242Z digest=sha256:0dd1705b0093044b1641ef011f450e63e6be04175956074d7bd8425bc8469426

Observation 2a227863-0077-490b-81b3-7e8ae235e776 · outbound

This paper cites Any-to-any generation via composable diffusion.NeurIPS, 36, 2024.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Any-to-any generation via composable diffusion.NeurIPS, 36, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.798707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.798707Z digest=sha256:84eba6e765318461d90ac2d86ff824075709b3591fe0cc1e9f41fea535430af2

Observation de4c8042-f473-4ec0-be7e-fa036b0c4a4a · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.864139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.864139Z digest=sha256:4a1bbbf8698c1c8698c779bdbdcaad5c3fe7da2f0e6bb5ba3995708d31e9b049

Observation bc8082fa-ff6c-40a4-86de-58673450e2ed · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:25.933048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:25.933048Z digest=sha256:fd1ccc0cec23c9bd473963fcc912aebd32fb48d367fea9d38412faf1c62f4e05

Observation c9438d87-f325-472f-852a-a530122fe4ab · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning LLaMA: Open and Efficient Foundation Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.025472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.025472Z digest=sha256:9fc787625dec1eefe4c771da4883f35f746a039b50d64d1cb35f98dbae8adcfd

Observation cbde4961-f1a8-460f-a1a2-23d466d71dd6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Emu3: Next-Token Prediction is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.121878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.121878Z digest=sha256:2416e64f732c25ddcb6a48d2a78f8f2f4c7b2c694c8bdbe067b2d9c76bdff4b1

Observation 3b1f686e-b87c-4b9a-a8fa-b6dced2d1622 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.245919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.245919Z digest=sha256:df640e44f783fb1e9458e375735df5c8caf0d84028c664a1f2eb43f1819c0b04

Observation e308dcb7-501d-44bf-9a0f-d4708432b45f · outbound

This paper cites Liquid: Language Models are Scalable and Unified Multi-modal Generators.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Liquid: Language Models are Scalable and Unified Multi-modal Generators

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.366317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.366317Z digest=sha256:8d7ac238fcb9d41caa6bfbaea99be88916b924ebd1c49bf48bb5c146cc711742

Observation 727c3c1d-52ce-4d27-bf43-8b46fbc2fb7f · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning NExT-GPT: Any-to-Any Multimodal LLM

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.460651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.460651Z digest=sha256:ffd2b6c8de77768eae4fc7c24a5b595b1a265acb347af6d427684d9caed30505

Observation fd41ded9-b835-43de-a77e-0011efa915fc · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.558833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.558833Z digest=sha256:4257f587b820019202ea83fa2e97ec8a485d55375434a5f1f94d805bd0052a3b

Observation 4c66c372-c2ee-43d4-9d91-34869edf7d61 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.689185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.689185Z digest=sha256:4cb7555b67b9756664d32f68a1dc07a8cbe459605aa38ecf5e42f6a527b77b43

Observation 17d2601d-6d50-4659-876f-f6998775a666 · outbound

This paper cites Qwen2.5-1M Technical Report.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Qwen2.5-1M Technical Report

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.847399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.847399Z digest=sha256:ac0eeea0c41321aaadc85963bef98ed9df6dcae65cc5fca032c5b143de966154

Observation 34296fd0-0dcc-4c8b-b655-9991d86c2266 · outbound

This paper cites Hermesflow: Seamlessly closing the gap in multimodal understanding and generation.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Hermesflow: Seamlessly closing the gap in multimodal understanding and generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:26.974393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:26.974393Z digest=sha256:d20a388627bdbef84a8ee6b3dd33afa48216064fc0f066c3f6ecd2c75427a3c5

Observation 598112ee-2156-417e-9f68-063ebb26669d · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert AGI

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:27.100136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:27.100136Z digest=sha256:66a2cda0524b734e8057ba065a6b692013561fc8b7be705052c3876f05f32502

Observation 1266f9e7-3c5b-4c78-baf3-94b4c1da3891 · outbound

This paper cites DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning DoraCycle: Domain-Oriented Adaptation of Unified Generative Model in Multimodal Cycles

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:53:27.611095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T12:53:27.204685Z digest=sha256:1f03eb7086030cb18d8d34839d44b6b08abb69d6642d1d0006fd3383b8e649c8

Observation b7ac9371-df22-4361-af2c-b5731326aa0e · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:27.275816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:27.275816Z digest=sha256:157a1c16542d3cf9472f98e3d460d3ea632c844ced88838ee345968754458316

Pith citing papers

Observation c63afbc4-b1d7-43af-ba5c-dc92b302b53b · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.089323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.089323Z digest=sha256:6e0dfef6e3bca477773ff82bd4968e46c575a4f23a7bbec104abc6338fffe73c

Observation 1b38343b-63f0-4da6-9d9e-c2958e2738fb · inbound

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models cites this paper.

SRUM: Fine-Grained Self-Rewarding for Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T09:54:26.172394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:54:26.172394Z digest=sha256:f08914effcbb80ce835306a0d28da8912d7a9688dd7bed17451c98555c65289b

Observation f65b3710-3e23-49d2-92cc-c271d97ff38f · inbound

LatentUMM: Dual Latent Alignment for Unified Multimodal Models cites this paper.

LatentUMM: Dual Latent Alignment for Unified Multimodal Models UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:43:17.438705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-20T12:39:28.058049Z digest=sha256:34d696325cad780d956adcbbcdf5c15d40f2f3e72698e87b015cee785a25ba83

Observation d3fe3d45-ba2a-4eb6-98ba-c8dccb1f69d1 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 195

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.711658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:b0ec77fc52e24bb9238104d2f9d819c02151c439098254b38f34a43a84c8050e

Observation a4bd0c27-f7ff-4223-9366-d8b551a7fab2 · inbound

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards cites this paper.

Ask, Solve, Generate: Self-Evolving Unified Multimodal Understanding and Generation via Self-Consistency Rewards UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-04T13:39:51.506849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T04:58:15.891214Z digest=sha256:19de7f52a6131d0a3c652a3742af3fd07128556c4ba22bd46b2575615d533e95

Observation 75fc4684-086d-4358-8cb7-ba8102369c62 · inbound

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning cites this paper.

Model Guides You How to Draw: Adaptive Visual Gating for Unified Multimodal Reasoning UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T01:07:06.861298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T01:07:06.861298Z digest=sha256:da2745b0208cecd9b881f6ae4fd94b81c5b84b4a9336e5fd65527fec9ec1f084

Observation a2b02145-2d74-4af7-b2cf-57e6af75cced · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:58.185926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:58.185926Z digest=sha256:54545d497620823ed94555c860c8256299ee036be7950c4a1e018dda84514609