Pith. sign in

Paper Citation Record · LEDGER

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

As of 21 August 2026, this Paper Citation Record lists 95 of 95 outbound references and 4 inbound Pith citation observations for arXiv:2505.14682.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14682 v1

Coverage vector

measured 95 of 95 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:34:55.930004Z

measured 99 of 99 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:05:33.892978Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T21:37:25.273563Z

Reference resolution

95 of 95 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved86
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 52f83fea-7511-4ab7-a0ec-43ee187f46b5 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.512594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.512594Z digest=sha256:c78a79e32917d5b22715e78af9eba883882a339747a3ad55662036804e210b87

Observation 3ccdf867-2c99-446c-9a6c-d8ebb6f0135a · outbound

This paper cites Flamingo: a visual language model for few-shot learning.NeurIPS, 2022.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Flamingo: a visual language model for few-shot learning.NeurIPS, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.538623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.538623Z digest=sha256:1f3db799b78412108a0d20cc99f601fe9a6425b85b1374c50b68206d0a8f0be6

Observation a29d6707-0930-4c4f-91ba-124d843d6482 · outbound

This paper cites Qwen2.5-VL Technical Report.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.573669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.573669Z digest=sha256:61537b01ab74cb19177debcf5c01e325c99994b8c497f4ce6d8e6765f545ec5e

Observation 723682f7-77e5-40ed-95d9-d0776ccbd48f · outbound

This paper cites Improving image generation with better captions.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Improving image generation with better captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.593742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.593742Z digest=sha256:22b16de6caafed19350e3c79e1b77403d5e4e750647cf6ebc4bbc4e0cd9bba76

Observation 1dd1ae7a-d72b-4dc4-85c2-b208561c3573 · outbound

This paper cites Maskgit: Masked generative image transformer.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Maskgit: Masked generative image transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.624294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.624294Z digest=sha256:8ae9ceaf177ae64af7e02f2e3c357a7d354325d7f60dcad36709ff6c80ff0896

Observation 41889b5a-3469-4d56-bcc6-237abb7e73b1 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.667370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.667370Z digest=sha256:9a7dce7f675478803fc70981d3dfc465f27d73f09b63926e3664b003b2944490

Observation fd00a909-2791-402a-9890-700fe45a6f9e · outbound

This paper cites Sets: Leveraging self-verification and self-correction for improved test-time scaling.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sets: Leveraging self-verification and self-correction for improved test-time scaling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.701551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.701551Z digest=sha256:676431ce68b9cbf913f20c27a9b8cbb71aa26b2f9819d8877c599a7549103c65

Observation b20ee523-5bdb-4ed2-b2b0-729660110fbe · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.737513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.737513Z digest=sha256:ee0c67bc4db3c2ebffec08cf1045164c7eda3b1b135d37502e5bd2ca422a99c7

Observation bb32b126-d539-4364-8e45-ee9b1ffd6482 · outbound

This paper cites DeepSeek-V3 Technical Report.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation DeepSeek-V3 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.773916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.773916Z digest=sha256:ea89fb71bc853397e7c2ebd36ecbde5f14b8adbe11fad5fb422782a50433b56a

Observation 55a7a45d-90ed-44dd-8235-3e0642eced0b · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.819499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.819499Z digest=sha256:f5113d4ccfa4cefc67187dd47abfd270b3e5ae6485b4e5ec28b3de441fdd7452

Observation c67c16d5-3ab7-46a2-917c-ff48f0d0e502 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.861983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.861983Z digest=sha256:0c20a95434cf9251238b2d3cefabb2cf2d0528d83f3d1bcd6000a03630f8cae1

Observation 3859a696-5813-410b-8b04-c23ae7268c3f · outbound

This paper cites Taming transformers for high-resolution image synthesis.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Taming transformers for high-resolution image synthesis

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.909519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.909519Z digest=sha256:5a8f09b40cdd69634ded43832750b825bb7bb65ddf6fcfdf1efea16afa12c971

Observation bb2098e1-9e53-41ce-a64b-85e74d3182cd · outbound

This paper cites GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation GoT: Unleashing Reasoning Capability of Multimodal Large Language Model for Visual Generation and Editing

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.941000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.941000Z digest=sha256:6712974d3170a74433bcef3ea6b544a5f8d0d9f32bcf461686c3bd12f7679677

Observation 26ad3188-5404-46a4-89a1-afad38099a2e · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:52.996397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:52.996397Z digest=sha256:f606959c542d514dcdaf1aedac638b2e67a5a3019a8917197dc05556a2389430

Observation f5efc5f2-9304-4718-aec1-1a86e142a83d · outbound

This paper cites Geneval: An object-focused frame- work for evaluating text-to-image alignment.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Geneval: An object-focused frame- work for evaluating text-to-image alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.003939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.003939Z digest=sha256:59d9a80abdc65faa547db18714d0fff2c3810f2b7065a20e57c9f6d288f5957b

Observation e2e84395-3dc7-48de-beab-b1a69a24e420 · outbound

This paper cites https://x.ai/news/grok-1.5v, 2024.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation https://x.ai/news/grok-1.5v, 2024

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.021250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.021250Z digest=sha256:958618cfb0c0c306b0b09eed2c00247f0b08b82be31e17ea4dfec8d4c4bcd735

Observation 783a8d68-ade3-4a30-ad42-6f0486bb4721 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.085420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.085420Z digest=sha256:b6bc9f80c3187ee776122746cf2d5d681f1c8797fb8fa57351a46e20c5d24dd6

Observation 1fc057b5-f3db-4e60-b51b-77bf28e69907 · outbound

This paper cites Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Can We Generate Images with CoT? Let's Verify and Reinforce Image Generation Step by Step

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.137002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.137002Z digest=sha256:134bbfdbfd8a22cc1a12fc560747e72c6c8e86b8ccd3889e8f5acad6898366a8

Observation 9ae4cd01-79ad-4ead-9d2a-ada7e66c244a · outbound

This paper cites Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Infinity: Scaling Bitwise AutoRegressive Modeling for High-Resolution Image Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.201462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.201462Z digest=sha256:a856522a6482a6a47674bf89fa1ea6dc296c4cf8b8e7202289b46b8c7ddc306f

Observation 53f55328-9d90-4735-b33c-25db97047411 · outbound

This paper cites Classifier-Free Diffusion Guidance.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Classifier-Free Diffusion Guidance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.259781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.259781Z digest=sha256:5d6afc9a43125ede90c1285ba2d3c87759236fffe45bb9731c3458367002cad8

Observation 0c8de696-4634-4b05-8551-d6ca78412285 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.306886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.306886Z digest=sha256:30578bccc785216736b04068b636b09254428a4700bb5ffdb23969a3a014d7f4

Observation 02d160ea-3ef4-4150-b529-792935c7a446 · outbound

This paper cites Efficient Test-Time Scaling via Self-Calibration.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Efficient Test-Time Scaling via Self-Calibration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.349405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.349405Z digest=sha256:ca7f659c0f0dbf09c0cce3413fa0df57d5ed81cf47afe8e6fb5539fd3b6fecc9

Observation d7a48737-5dac-4fcc-af25-5e4dacd87703 · outbound

This paper cites T2i-compbench: A com- prehensive benchmark for open-world compositional text-to-image generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation T2i-compbench: A com- prehensive benchmark for open-world compositional text-to-image generation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.404595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.404595Z digest=sha256:e7819de1768a9d394dfb7750ec56e5ab4ff083c501402870f0aa040ce3c4df58

Observation 4168fd3d-96c2-4cf2-a42f-8dcaab03c391 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.466962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.466962Z digest=sha256:224c9787d1bcde527fd04a1d42b790391fcdfdd356591a35e98787a52d214a9b

Observation 94a9d7f0-ab74-4b31-8f27-4d6f0e347e4a · outbound

This paper cites Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback.NeurIPS, 2024.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unpacking dpo and ppo: Disentangling best practices for learning from preference feedback.NeurIPS, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.509431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.509431Z digest=sha256:e5eee9675225644fe32d4390cc0d45ebe24bde90f7918c5d37f5d3a70f5a7b72

Observation 90b0f1c6-29d5-4497-9314-3dc8be323414 · outbound

This paper cites text-to-image-2m: A high-quality, diverse text-to-image training dataset.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation text-to-image-2m: A high-quality, diverse text-to-image training dataset

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.546081Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.546081Z digest=sha256:044936adddfe2b2c1d5efeed9c57a5708fb844879fc342ae0ee17912e75cd0f3

Observation 31234611-dab3-4a1d-a3d0-a54352ddb84b · outbound

This paper cites UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation UniToken: Harmonizing Multimodal Understanding and Generation through Unified Visual Encoding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.601708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.601708Z digest=sha256:24f4b66097c3db3a12206218c87df87935d74e83bbca26128b2666890622f55a

Observation 88929cfb-eeee-407d-86d5-fbb88bd7afc1 · outbound

This paper cites A diagram is worth a dozen images.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation A diagram is worth a dozen images

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.658834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.658834Z digest=sha256:e38045bb423a0dc9256bdce3a061a1289a21bbe61a43bdc5c8ff3874bf445585

Observation 982002e0-8e72-4afb-9316-9fb10a7b138f · outbound

This paper cites Segment anything.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Segment anything

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.714539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.714539Z digest=sha256:bf91b72bc5dee069a85f3b5afe0c2250f87233f4d1600444b4cc180b928d2832

Observation d56196af-47df-4c75-96f1-e351b73c2c68 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-OneVision: Easy Visual Task Transfer

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.781304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.781304Z digest=sha256:f688b212064f05c1061afad0aec9aa20bc6b3598e52945cc91e8d6e600379108

Observation d629b1ae-9e12-4c2b-84e2-de35dadc566c · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.828874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.828874Z digest=sha256:4bc9fa73967890cb9610d734422adb29400ec2a355e224e3cd8fcb39a470eaf7

Observation dbd58054-26e3-43d4-9727-aa8e40f15f80 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.874923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.874923Z digest=sha256:6525129e5e56336148ee3705d41cbf4aa9885b44a103632565d62120106b8f08

Observation 04f4078a-e59a-4243-a4d3-993af8aa5195 · outbound

This paper cites Evaluating object hallucination in large vision-language models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Evaluating object hallucination in large vision-language models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.920486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.920486Z digest=sha256:18401e117532b070b626de665ef2e183a17d36195678a88f8372cea0b3e32adf

Observation 9bbb5788-a59c-4612-a250-14e6e2bce774 · outbound

This paper cites Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hunyuan-DiT: A Powerful Multi-Resolution Diffusion Transformer with Fine-Grained Chinese Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:53.968678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:53.968678Z digest=sha256:b016f0aaf1015c316d97c84c70f17535e8cdacbb03609c52f2f3ec915c97cf15

Observation ca4ad6c1-b257-4436-a767-6af5ff7bb5d2 · outbound

This paper cites ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ImageGen-CoT: Enhancing Text-to-Image In-context Learning with Chain-of-Thought Reasoning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.013136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.013136Z digest=sha256:c7313f75b4bf7ec4d69411a3abffdbb7e3d1bd161c6f3b1342b4485ba4edb400

Observation 45e4b063-c515-41d7-a93e-82ce1dfa65cd · outbound

This paper cites Improved baselines with visual instruction tuning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Improved baselines with visual instruction tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.047031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.047031Z digest=sha256:c4613911894aa0d8be2eec3c62adbd20cf7f6bbd0f5d99f049093263f4f6e4c8

Observation 30077521-d764-4357-aae6-cbc02cabcb78 · outbound

This paper cites Visual instruction tuning.NeurIPS, 2023.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Visual instruction tuning.NeurIPS, 2023

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.097101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.097101Z digest=sha256:eb66c844d5562839c582e23629ae5073aeb81297cac5e6c8875a2e8eb8ac21af

Observation b6368733-9fef-477b-b20f-a6c9b9c6ef7a · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation NVILA: Efficient Frontier Visual Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.126661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.126661Z digest=sha256:4b5265bc982ff8acbd69df21bffe8a3adcb8113220c83a2a4cd19c942c511e14

Observation e584414c-19b1-4a33-9fa2-6928c587e898 · outbound

This paper cites Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Oryx MLLM: On-demand spatial-temporal understanding at arbitrary resolution.ICLR, 2025

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.828679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:54.151515Z digest=sha256:d93ba3705c6d5ae49eacf56d66d116740fff8e1a1aa9936159148fff28049469

Observation 3bb90c00-50f0-47ca-b7e3-50a3fd2f7d87 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.170183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.170183Z digest=sha256:42821802f285893c0651f76dff6247518e950e19d74878458d52dfa89eb6a109

Observation a974054f-a8f6-4b30-9fba-5ae051fc1008 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.204865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.204865Z digest=sha256:e9a6b8c3a5b6da8985ba1670778254367062c2f10b3538fefea5d4ad73e24d2b

Observation 3a559694-8a17-472c-8603-2ba29b57c09e · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.231257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.231257Z digest=sha256:7c632d00a6b405bc39f4a78a67ba9a2042038bdb3b56bb6aaf92ad5b5673fd06

Observation e0d2967d-0f9c-46f5-8a8d-af0fee297e29 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.260170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.260170Z digest=sha256:86ce324981a5c09032b16b268da7fa49691dbaaf79cae3f0da873b69d14ccae0

Observation b9e736e6-6583-4c76-9921-b5aa0b8a23bf · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.344523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.344523Z digest=sha256:56ce71c154936ab1f75453f38237df9d1f4965817ce0c580d5a9bbddf81f18c9

Observation 67ee5d29-e318-4817-8448-150ef028d871 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Finite Scalar Quantization: VQ-VAE Made Simple

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.404034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.404034Z digest=sha256:da3a0e5eaff57ea4b7a901668445350a6e8fa65123c8df0ac5786b350b9809ea

Observation 9c6223bf-0d29-4da6-9e12-a8d5af6071e9 · outbound

This paper cites 4M: Massively multimodal masked modeling.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation 4M: Massively multimodal masked modeling

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.804787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:54.466396Z digest=sha256:c0feeb15911d4f665247d810a20f2884e147252d91e4605e94c7ef8e3c5f5b59

Observation 3a0b26d5-444a-44ca-8cb2-4ad988125365 · outbound

This paper cites Gpt-4o, 2024.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Gpt-4o, 2024

Reference 48

Resolution
parse uncertain
no resolver link, observed 2026-08-07T15:34:54.551750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.551750Z digest=sha256:a6cd427f358f98384f4141df43602e9899dd2d64b2834217fc1a789d25ad17f0

Observation 72ed0b6e-5140-482c-bf39-a60864f2e53a · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.618377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.618377Z digest=sha256:f0582bf526b93653d726ff22994370ce802aabef12d05125b85c6456ce66ebb1

Observation 457042c4-d2c9-4fdd-a391-87a1a1d0f3a2 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.687530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.687530Z digest=sha256:3af2c51fda56b98d226837e549ac9542b68abc6a2f756c6a6f3c1c4ab258e89b

Observation 078059fc-f144-40af-b551-fc62dd3a4209 · outbound

This paper cites TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation TokenFlow: Unified Image Tokenizer for Multimodal Understanding and Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.732431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.732431Z digest=sha256:9600ec3bc4a223195514496b5cdd22c345ff650e33d5243e2833422b9a5784db

Observation f242a8cd-9bbc-4ae8-9eb6-7544a417f009 · outbound

This paper cites Learning transferable visual models from natural language supervision.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Learning transferable visual models from natural language supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:54.938943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:54.938943Z digest=sha256:bceb618d3e0329a701125099a385646376c7ee0f79ec53a0ab1a7d538efd2d64

Observation e6a5ed8e-c7f1-4399-8062-165a553d84e0 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Direct preference optimization: Your language model is secretly a reward model

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.046840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.046840Z digest=sha256:fea01c2e06e822012143015068203f287d08af5731127d6ac71bfd269d6bb9ab

Observation 4e76550e-fba5-4a35-b15a-a9b53de0db40 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.250725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.250725Z digest=sha256:1df0c77bed97fe8bd968bd81279e0b76b495097c6a1f9a9a5522babcc1c7f11b

Observation 34142080-8a96-4211-8d6d-048ac101e8a3 · outbound

This paper cites ImageNet-21K Pretraining for the Masses.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ImageNet-21K Pretraining for the Masses

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.362129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.362129Z digest=sha256:eca41736707c69264a63f4c73912ad7156d42bad82f12d2f5f89b412a60107b6

Observation 37da172c-6c94-4ad1-85e4-c40fd9e49d2e · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.466549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.466549Z digest=sha256:0a53ddc500708e635826e0ba38cc46c1c9077f45535a259f1b515938b7eee8be

Observation b34b7e8a-a3e0-42ba-9eb0-12d860fa1324 · outbound

This paper cites Journeydb: A benchmark for generative image understanding.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Journeydb: A benchmark for generative image understanding

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.758086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.654654Z digest=sha256:d1592c7a1d1194d1456bb2d7ae5e21c85ed414e03add5875c05b35dfced7b01e

Observation 20e719e6-c461-449c-b2ac-866aa70f6eff · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.760316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.760316Z digest=sha256:93df41421ca30cef77bbc7af1bb5228408c48bef8843baff4c6955ef920544a6

Observation 83b385a8-a10e-46ef-ab1a-eea6ce7f4a7e · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.770576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.770576Z digest=sha256:aa6da450148cff66445537b7d18c7701321f750ed6c403a49d76e9bc588412f6

Observation bd113272-fccd-4aa4-a22e-de30ec086e59 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.778208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.778208Z digest=sha256:cf6c2bad769d7dedb6da66718dbc24431d5dbf76195a912aa1fc1603433affc6

Observation d7179324-bc70-4faf-99fe-2111eb1aa559 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.747202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.784427Z digest=sha256:73c365dd540f4b928a9df7ad106324a652fb70a576e80af9c3e986b6ccd1872f

Observation 4bbe62b5-fde3-4c19-b30d-e018fee5bd08 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.788642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.788642Z digest=sha256:d1c86568393ef3293d6eb7650296014e8996ec0e205f6c49816ba397eed51eb1

Observation afa9970b-f0fa-494c-bb45-39a1cf410247 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaMA: Open and Efficient Foundation Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.792791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.792791Z digest=sha256:49c65305757433e6daf7f28955f3bad708ccf2941115d586744b23210af6873f

Observation ad0c2970-0cfd-4b2b-abe8-f93d345c7370 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.797059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.797059Z digest=sha256:01b24f8bfd5dd6589fc9ac8a2e10553fdae6d0ddeeee2209a927b49fad7a3382

Observation 846badcb-5f5b-47e3-8b07-2746accc0d45 · outbound

This paper cites Neural discrete representation learning.NeurIPS, 2017.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Neural discrete representation learning.NeurIPS, 2017

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.801017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.801017Z digest=sha256:2c612ef97cccd86df67713d9a8834678b020cdda11882aa177e8b7419bd58760

Observation c7126fe6-580b-4126-8f5d-7bf49f15da5a · outbound

This paper cites ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation ILLUME: Illuminating Your LLMs to See, Draw, and Self-Enhance

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.804718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.804718Z digest=sha256:eaed0827528d1acf63d709c16a0b3509be52ef8972d994a4b0e3d82b5547e1b9

Observation 0dd2d8fc-0dd2-4746-9008-f3120daa65f6 · outbound

This paper cites SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SimpleAR: Pushing the Frontier of Autoregressive Visual Generation through Pretraining, SFT, and RL

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.808400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.808400Z digest=sha256:4782f5f9e82a2e0677499fe9244555f366d3bf7c40be1ce7a8454895e1cf72e2

Observation 797fd783-9290-4779-a178-3ce9d99374b9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.812726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.812726Z digest=sha256:9df5fac0a72c8fce3c905fdfd84d29d048f0bf87a053b74f11347849a6e56be8

Observation e50f3828-5425-42c2-acaf-68bb2490f3fa · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.815984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.815984Z digest=sha256:b8f16634bf6ab79e19bc2da67be0821843d1ada633047134cc4edd2b775b2227

Observation e1b25499-ef54-4575-90e3-e45f4e117118 · outbound

This paper cites VisualPRM: An Effective Process Reward Model for Multimodal Reasoning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VisualPRM: An Effective Process Reward Model for Multimodal Reasoning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.820224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.820224Z digest=sha256:eb7a6cc1eee8eb36f6f566f0277ec5d475fef6fb83f4d967cbff3ea660002e81

Observation ea95bff2-5f5e-4351-8b28-03f390050f15 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Emu3: Next-Token Prediction is All You Need

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.824650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.824650Z digest=sha256:db74d3e5a5727993ff4c8689bb0c9ac40ffbb3f612d4c0fa5777f9ce598d3bfd

Observation 0cc76da2-e4ec-43f6-9233-39a53873b3f8 · outbound

This paper cites Mint: Multi-modal chain of thought in unified generative models for enhanced image generation.arXiv:2503.01298, 2025.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mint: Multi-modal chain of thought in unified generative models for enhanced image generation.arXiv:2503.01298, 2025

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.828840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.828840Z digest=sha256:e5d8cc15dd709354a739adebb3def8ffe0eb0e16e9cf8658a69c2807a48e63c5

Observation 2d61d451-a3cc-4c58-b64b-c87b07c62c9e · outbound

This paper cites Large language models are better reasoners with self-verification.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Large language models are better reasoners with self-verification

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.730638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.832641Z digest=sha256:ed5f95d48f878f347e147944be48b5a1decfeb38cd7d5ba734cd06288ababa29

Observation 728e9b3f-4e14-4fce-8b4b-57d53df7a461 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.836812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.836812Z digest=sha256:cc2240ca479029fbe452517b9706b7e1f170a4c8090c06b48ee61b202d08dfb4

Observation 0d0cf564-ffd4-4eb8-97aa-2616e011c4d8 · outbound

This paper cites Vila-u: a unified foundation model integrating visual understanding and generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Vila-u: a unified foundation model integrating visual understanding and generation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.720644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.840962Z digest=sha256:059312b331eeee24b75d26fe27670249aadfa611ce09eb8d4e8ed5e53fcead57

Observation e81342d4-62d6-4e24-a7e8-b08ab321e5d9 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.845056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.845056Z digest=sha256:7267184a1e1750d8724d2cf757e0290457af395964d6c1783ffbb64d137ecfbd

Observation 9c0703a0-d95f-4e4d-a78c-9d84baac40c0 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.848923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.848923Z digest=sha256:531eaf62d70daa8d1f8d06c20a5e16634067bad4d36ad9fb984a33eca149d42e

Observation 0121dee9-6d54-4503-89a9-2066e849809e · outbound

This paper cites SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA: A Strong Training-Free Baseline for Video Large Language Models

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.852874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.852874Z digest=sha256:4a0c7ec892fd26ee8c36b1a242bdea8f4c2b7e6fa759f6e282a45aa358c48682

Observation 765079fd-8673-4ee6-80de-35b19f652f4c · outbound

This paper cites SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation SlowFast-LLaVA-1.5: A Family of Token-Efficient Video Large Language Models for Long-Form Video Understanding

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.856714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.856714Z digest=sha256:c6081053b6a6bcb2e35cadb3d917f7994b4b28cd0673b72b00392bc45989c7ae

Observation c61bde09-8bb1-41a4-91b0-14fb9ca232ac · outbound

This paper cites Qwen2.5 Technical Report.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Qwen2.5 Technical Report

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.864539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.864539Z digest=sha256:e9c9650d8a46c06ca68672cf49438174250535af745990f5315739b0bd3591a0

Observation 5292e68c-835e-4792-9d06-564ade5c221a · outbound

This paper cites MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MMAR: Towards Lossless Multi-Modal Auto-Regressive Probabilistic Modeling

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.868975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.868975Z digest=sha256:05d4a557fc5bcec423e2229960a9a2d273492a22e80f2d89884279184a84b08c

Observation 799d2c91-9672-4fc5-8b41-02c3ef15c8cf · outbound

This paper cites Hermesflow: Seamlessly closing the gap in multimodal understanding and generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Hermesflow: Seamlessly closing the gap in multimodal understanding and generation

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.873451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.873451Z digest=sha256:85c98ad0bb83ef9039ae9df767551d13222933b61930aecbadec6277f9c204a1

Observation 5fb6049a-a965-430f-a92e-3813bd5b6a78 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.877402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.877402Z digest=sha256:c37a3b4b4066c7605779c5b951959f93adfb9b3a03dae15874a6b5e7c9f9f96f

Observation 4806fd98-19bc-45ea-9e7f-f0b3208da644 · outbound

This paper cites X-VILA: Cross-Modality Alignment for Large Language Model.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation X-VILA: Cross-Modality Alignment for Large Language Model

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.881448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.881448Z digest=sha256:c8499a96c5dd5445d772ec8ba92939db0e76eda067e1203d16edf99d6ae900a0

Observation c52337af-fc4b-4ce0-8662-eb2fe8548418 · outbound

This paper cites Language model beats diffusion-tokenizer is key to visual generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Language model beats diffusion-tokenizer is key to visual generation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.710212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.886007Z digest=sha256:f13daeb7f01e74a51528486716489e82cfaad68b1c2f3a88895243a162248bad

Observation 643bc847-d865-4585-adc9-b1dbe937335b · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.889892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.889892Z digest=sha256:7628d6fb61285ff472ebd3b9b7edee5872dc4322cad14a33247bef87102d771d

Observation 97a2c1ed-dfcb-4c1b-9e16-a9853daad063 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:34:56.697844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T15:34:55.894512Z digest=sha256:ed44b6701278762460edb321c1e9998f34d80e32d454e7d6e457f1c80ce8e2bd

Observation 69a51502-6e1c-465e-bf5c-450459cb3dd4 · outbound

This paper cites Sigmoid loss for language image pre-training.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Sigmoid loss for language image pre-training

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.898317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.898317Z digest=sha256:50429f0142d2246b8ef3d78ef9bb9123b7e3be302cd18b8fb0ee89afb062e911

Observation cc3c874b-a5c4-43a9-b2fd-aca20f4a9131 · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.902559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.902559Z digest=sha256:cbc234f94891e5fe4e475bc40a7d5556506377e1be9876c2f8c566069a184678

Observation 202f015a-ce24-49c9-9654-610a49bdaa9c · outbound

This paper cites MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation MM1.5: Methods, Analysis & Insights from Multimodal LLM Fine-tuning

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.906384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.906384Z digest=sha256:de6f2150ceeafd5d3678786707cdb8867646f4d6dd73b406efb0b3cb2acf3484

Observation 37f6e468-d31e-4e85-a9f6-ac9bc1835124 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.910000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.910000Z digest=sha256:f4cce2ea9663a85dc651894eeb996abc7464e86a18b0235f638b22dd8c0ec07c

Observation 53a839f8-3586-4ffc-af44-54354bb9b2c8 · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Image and Video Tokenization with Binary Spherical Quantization

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.913350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.913350Z digest=sha256:9964a36a250ce27072341044128df442e281cfcb2ddea7dde83a6a51760abf7f

Observation a8fddcf2-8c72-4318-ab50-bdb79916dc93 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.917512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.917512Z digest=sha256:fedec720c80388a55018f75b68ce9f9b6ec850741c63abcc7fcaea0da896f55d

Observation 0e74cc5c-5dd6-4746-92a7-02b407ac0979 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.921584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.921584Z digest=sha256:6e43eb03f1e07660723df5b86c9c7ac65b3245da8f441cf840fae395a590b6c8

Observation e4523b3a-873e-4504-abfe-4967ea71a8ad · outbound

This paper cites VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation VARGPT-v1.1: Improve Visual Autoregressive Large Unified Model via Iterative Instruction Tuning and Reinforcement Learning

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.925498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.925498Z digest=sha256:439fda716ddcd9d54d83c780c5f281539eba96f3c92dff4a08d346d5f32e8405

Observation 208b340b-771f-4eda-aa3a-b4fde362b2f0 · outbound

This paper cites Apollo: An Exploration of Video Understanding in Large Multimodal Models.

UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation Apollo: An Exploration of Video Understanding in Large Multimodal Models

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T15:34:55.930004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:34:55.930004Z digest=sha256:8202944df0145fccd15bf4d13ab61c3a80d580c383b684d8207e7b5bbd8a1d2b

Pith citing papers

Observation 7529171f-8a64-4c7a-bce7-ccd3c5170588 · inbound

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation cites this paper.

CreatiLayout: Siamese Multimodal Diffusion Transformer for Creative Layout-to-Image Generation UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T22:05:33.892978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T22:05:33.892978Z digest=sha256:6bac518d365aaf9c3fcf6745b3edde4a37ce447efc43677e270da3398d268251

Observation 3e6ede4e-781a-4854-8837-c1962009ed84 · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.663982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:8a14939290d7f59365423e936f19d6294c0a277bcc4e70dae7d63810a1fcd830

Observation 1fd8eb57-a07d-4e4e-8bcf-eb0a8e181796 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.154553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.154553Z digest=sha256:663a6c1898cdfe1cd28703424b9dccaed7e6284a1fd392588e150fe099cd462f

Observation 3c4b1489-358f-4687-aae1-d500560bbb39 · inbound

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning cites this paper.

Test-Time Scaling in Multimodal Foundation Models: A Comprehensive Survey of Generation and Reasoning UniGen: Enhanced Training & Test-Time Strategies for Unified Multimodal Understanding and Generation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:37:25.275014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-27T19:36:57.231932Z digest=sha256:d407d926639a1b51a2494b2778a5844ff8e554d966d9703b935c6f29a056e341