Pith. sign in

Paper Citation Record · LEDGER

Ovis-U1 Technical Report

As of 21 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 34 inbound Pith citation observations for arXiv:2506.23044.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23044 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:01:21.610269Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:31:05.911595Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-05T16:51:14.255860Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1f26940-a111-4509-9044-f93b68e15997 · outbound

This paper cites Qwen2.5-VL Technical Report.

Ovis-U1 Technical Report Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:18.890898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:18.890898Z digest=sha256:04e415142cfb78a893760d8371c679776fb62c97da47b34f72b962b349733590

Observation 1dee3b2c-2436-46d1-b713-42165f67ae27 · outbound

This paper cites Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction.

Ovis-U1 Technical Report Ming-Lite-Uni: Advancements in Unified Architecture for Natural Multimodal Interaction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.342279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.342279Z digest=sha256:e357c9c1b3ecf3357ed06b0a2968e45ca921dc45b479e0442ecdf9e10decb625

Observation bfebae62-49c6-4f63-9b2a-40380c2b9677 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Ovis-U1 Technical Report HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.699701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.699701Z digest=sha256:3149da069e50ce538383a3a01d0a859cee90707882482abb9e9aa821b1e7cfe8

Observation a7f30d90-fd70-4685-adca-9ceaab2d380d · outbound

This paper cites Generating multi-image synthetic data for text-to-image customization.

Ovis-U1 Technical Report Generating multi-image synthetic data for text-to-image customization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.759372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.759372Z digest=sha256:c70e2ad308d8f454a13c3b64e7f0a9db3eba3581d7b67f714f6472953e3810e6

Observation ff295d1e-38c2-4eb2-8bb2-3530cd5c1afd · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Ovis-U1 Technical Report FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.838068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.838068Z digest=sha256:c79c38c370554c6731d00f12e147e267637494368489b63aff1c7cd37dc42885

Observation b1c27ebc-14f2-4f97-ba9e-1b17c909838e · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Ovis-U1 Technical Report UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.948503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.948503Z digest=sha256:f4a74a819ad0d76f26ebc42d8310c4f269ef544bff2b27f4b521bd41d6c668c2

Observation de2cc211-29d7-4ada-92a8-c56fefde7204 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Ovis-U1 Technical Report Step1X-Edit: A Practical Framework for General Image Editing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.117456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.117456Z digest=sha256:e22a2656efd8be3d5d086bce7174a5e5b740cdbbdba286718ec46e4383d7cd36

Observation 6966aa3f-f1fa-4f2b-9869-9325d6035023 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Ovis-U1 Technical Report MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.236642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.236642Z digest=sha256:e483da767dba6e30702cc7a9c6a15308c4cefecd5eb359e23e8ef5c182c56ba9

Observation 2aad56e1-aa05-4038-b813-07d979938afa · outbound

This paper cites Ovis: Structural Embedding Alignment for Multimodal Large Language Model.

Ovis-U1 Technical Report Ovis: Structural Embedding Alignment for Multimodal Large Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.362999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.362999Z digest=sha256:d8a6ff5e797365c2486cef370d6f116375c97c84b0e9773645e00fafae66b128

Observation a4c83747-2718-486e-b673-aa6780cd42c5 · outbound

This paper cites Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models.

Ovis-U1 Technical Report Exploring the Role of Large Language Models in Prompt Encoding for Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.485201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.485201Z digest=sha256:afda09e0426d5d45e422b0bbab7e5a7fef2e4d9ca23f9c28501ee0687ae0766b

Observation 832f4d8a-2664-4068-8465-4905474e6b9b · outbound

This paper cites OminiControl: Minimal and Universal Control for Diffusion Transformer.

Ovis-U1 Technical Report OminiControl: Minimal and Universal Control for Diffusion Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.775908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.775908Z digest=sha256:140f3ccbf1036ce379a92ee0467c0fc6c461fb535c2e739f748a4801e2b6e555

Observation 19406f29-e030-43b2-a19e-d4750f808bb6 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Ovis-U1 Technical Report Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.910350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.910350Z digest=sha256:e3f063132ed7e775c822879272c3c47280ec1b152d43c42985a030b8f129a433

Observation df0d8873-5fd2-4ce5-9671-a01b95f3a8e0 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Ovis-U1 Technical Report OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.967384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.967384Z digest=sha256:3da0a2a4d0eb1520c39dbafbf2dfba8a39f1b0b3dfd28aa27fcccc0e81e29ae5

Observation 8cf80178-8f4f-4ce6-b773-27d2b2e069ee · outbound

This paper cites Qwen2.5 Technical Report.

Ovis-U1 Technical Report Qwen2.5 Technical Report

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.047190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.047190Z digest=sha256:8c62bbc88abbf1c2f5965fd0f794fc94adf3e9eab4437182ea6707b74f315d6e

Observation f6b919d7-06ec-46af-943d-2ae1d285fe96 · outbound

This paper cites Qwen3 Technical Report.

Ovis-U1 Technical Report Qwen3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.141916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.141916Z digest=sha256:087351aa337c95e1ae1b60cb9360fff21e0794ea16804d7f25a10637093ac7ec

Observation 34755a93-0a9f-4799-a54b-94c234da339c · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Ovis-U1 Technical Report ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.241209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.241209Z digest=sha256:eeda9024013ce1a0ade590671c01898c69e10e6150eba3a7d82e11fe6d3d3b1f

Observation 37e62b70-b871-4fba-8415-04f74c3b61b2 · outbound

This paper cites Unified multimodal understanding and generation models: Advances, challenges, and opportunities.

Ovis-U1 Technical Report Unified multimodal understanding and generation models: Advances, challenges, and opportunities

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.423644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.423644Z digest=sha256:954adfa99018945ad452d242ca8266a22db1e71569402114f3996380724f5940

Observation 27412aec-328a-47fb-9fbc-581d57699664 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Ovis-U1 Technical Report InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:21.610269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:21.610269Z digest=sha256:b4adb18f8b3bd8be1a13b8b763c9137df71a531c4b4bc45ae5fbc536905e22cc

Observation 1f4da6bf-77e9-4a35-829a-8cbaf1ad1744 · outbound

This paper cites SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing.

Ovis-U1 Technical Report SEED-Data-Edit Technical Report: A Hybrid Dataset for Instructional Image Editing

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.198105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.198105Z digest=sha256:3af9c58d6ab8f6838f05579a54f1c360719091543191b4bfd2d7fc2214dffebd

Observation b14cbdff-083d-445e-8ddb-7ce25a85e739 · outbound

This paper cites UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild.

Ovis-U1 Technical Report UniControl: A Unified Diffusion Model for Controllable Visual Generation In the Wild

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:20.596851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:20.596851Z digest=sha256:81fd10cb5fa9b8da960aac0de45d3b57253d3aa4a6b0adb0414ac47d0af2db3b

Observation e0cc0b9b-74db-4cd5-b9c9-2db2a0241973 · outbound

This paper cites StyleBooth: Image Style Editing with Multimodal Instruction.

Ovis-U1 Technical Report StyleBooth: Image Style Editing with Multimodal Instruction

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.476640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.476640Z digest=sha256:e7de91610511169a849d56b2897637764b4b597df170a8d91a92de0b5f394d11

Observation 7fffe3c9-409d-4de7-ae26-296c0c759451 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Ovis-U1 Technical Report ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.522547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.522547Z digest=sha256:35dd3e508170c483c27338fc117d8d5fd1e0169fabdfd5d1da89947060c9a474

Observation c54a4482-8029-4f7e-8993-d4e00c5153ac · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Ovis-U1 Technical Report BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:18.962171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:18.962171Z digest=sha256:5e89da04bca577e0f733306f81b0185a5362765e52f71c5ffb53a3ae0316f58c

Observation 99b5d5ba-653a-4696-b463-c2707af7e4f3 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Ovis-U1 Technical Report Emerging Properties in Unified Multimodal Pretraining

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.045935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.045935Z digest=sha256:e62076cf0cf9cb7fcc2c8d1849d91c37664a3e4888fc1576f0532411523d6f67

Observation 0d02e96e-c3c7-4bd1-86f5-1a7e2fc38f73 · outbound

This paper cites A diagram is worth a dozen images.

Ovis-U1 Technical Report A diagram is worth a dozen images

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:01:22.180016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T22:01:19.593095Z digest=sha256:d22fe202a162a0bc7c0fb798db2d8c4111a1b868dcd25edbd6f54436812a1deb

Observation cda295d1-e7e2-4768-962b-fe867de61575 · outbound

This paper cites Scalable Vision Language Model Training via High Quality Data Curation.

Ovis-U1 Technical Report Scalable Vision Language Model Training via High Quality Data Curation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T22:01:19.113047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:01:19.113047Z digest=sha256:99cc41b3cb45f4a657859d52ecf5b18925e31d5179060e204684d9243a878b69

Pith citing papers

Observation 9c05e8fc-c730-486f-bdc0-f7ccc969fa2d · inbound

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again cites this paper.

X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again Ovis-U1 Technical Report

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T12:10:08.510929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:10:08.510929Z digest=sha256:a53ded19e7f04d220b4f050121fad42a9a90b1a18ae5d6b45f896d49e4bb2f79

Observation 68f62def-fac4-4a54-b902-6a6326680544 · inbound

Reconstruction Alignment Improves Unified Multimodal Models cites this paper.

Reconstruction Alignment Improves Unified Multimodal Models Ovis-U1 Technical Report

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-04T22:36:08.177277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T22:36:08.177277Z digest=sha256:a2f7b0e30fe1785c9066872e342ff72d36ba63f865e555daa4ae86f187a5b5e3

Observation e4e1981a-d8f5-4e2e-a892-9b6253aac7b7 · inbound

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples cites this paper.

Few-Shot Synthetic Image Attribution: Identifying Unseen Generators with Limited Samples Ovis-U1 Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:09.708983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:44:09.708983Z digest=sha256:560d187b0b34ffb0496031e3ea99907aa0662bbb2bd7ee1e7c0ece978181f404

Observation 388417b9-e942-4903-b2e5-7420ab00a381 · inbound

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs cites this paper.

InfoTok: Information-Theoretic Regularization for Capacity-Constrained Shared Visual Tokenization in Unified MLLMs Ovis-U1 Technical Report

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:10:45.413995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T08:09:19.209759Z digest=sha256:08bc0522b8768bb8b45c4e4301ff71b675efd90ccc8d02267f5b41659f9fafa1

Observation c63b1aef-16f2-4797-ac9a-c00c190a6fbf · inbound

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks cites this paper.

PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks Ovis-U1 Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:00:42.767279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T07:00:18.807922Z digest=sha256:4b75e78bf35fa48cfdfedf92b9c7cdb54ef5a2ef6c11eb31eacce7b376e00314

Observation c0b3b0ae-925d-4c11-a52b-13e69920a4fd · inbound

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization cites this paper.

OmniFysics: Towards Physical Intelligence Evolution via Omni-Modal Signal Processing and Network Optimization Ovis-U1 Technical Report

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:10:43.051870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-16T07:09:46.254851Z digest=sha256:aabccda4137d1fcd0c1e62bdee9d2a4a7403685ced99a0ed1ca2f11256d8d0fb

Observation 72272e35-bc7a-4e93-98b7-74e476b2700d · inbound

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy cites this paper.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy Ovis-U1 Technical Report

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:b38a00b68c9e924d832667ee51b6dee969f68ac2adf1da15138eb06fbad90696

Observation f735d185-2c74-49ba-98d9-a19ea1605fe2 · inbound

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking cites this paper.

TRACE: High-Fidelity 3D Scene Editing via Tangible Reconstruction and Geometry-Aligned Contextual Video Masking Ovis-U1 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-03T02:31:04.075503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:31:04.075503Z digest=sha256:a7f9d0ead8694a01f0e4e372a2aa979a501e8518821c9172b28ad02c2597ae00

Observation 2c9f300d-95f8-4824-a1c9-a6074134a83c · inbound

Training-Free Image Editing with Visual Context Integration and Concept Alignment cites this paper.

Training-Free Image Editing with Visual Context Integration and Concept Alignment Ovis-U1 Technical Report

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:05:47.728342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T20:17:34.383852Z digest=sha256:1b72d212f96080fc4632b8f0d9680e87d679e2b4f10137c9fd5c604aa5761f36

Observation 111eae68-f7db-4cd1-a53d-15684b85ba06 · inbound

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment cites this paper.

Learning Preference-Based Objectives from Clinical Narratives for Dynamic Sepsis Treatment Ovis-U1 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T22:22:06.385856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:22:06.385856Z digest=sha256:55221574cd427d2a1e9715d38f29619f3688bbe868f2a207b34c12567222982f

Observation 769cea6d-9b80-4643-bc85-90fff50cd8ce · inbound

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models cites this paper.

Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Ovis-U1 Technical Report

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:50:59.247101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T16:27:14.491492Z digest=sha256:13d225226b22730a339461fe2aacc468cf921ff3bc6658c04d82fbf24f05d298

Observation 58d72faf-f4a2-43e7-b1fa-44eb7d8d53e1 · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Ovis-U1 Technical Report

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:18.987516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T06:03:06.920592Z digest=sha256:1a32c34ff57a8ff9d14fd851385162e7360ccb0b7b7e5b6548fee0201363a33f

Observation 6141e825-fa89-48f2-9b39-1e23e0800b4f · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Ovis-U1 Technical Report

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:26:29.544348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T02:52:58.232585Z digest=sha256:f964346b2a6e7e02d47903bdc0801269dc2dea8fff1470b0d914794c5ed1ea9c

Observation 1337908f-8e02-472f-b4a9-bad8915ae98a · inbound

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models cites this paper.

UniGeo: Unifying Geometric Guidance for Camera-Controllable Image Editing via Video Models Ovis-U1 Technical Report

Reference 68

Resolution
metadata mismatch
local_arxiv, observed 2026-07-05T16:51:14.258187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-07-05T16:47:32.853010Z digest=sha256:4913666d3d1ea4101abe642390f8ad2807136167388d4364c1dd185d53dd5d9a

Observation 975c441b-446c-404d-b0be-d206e36050e7 · inbound

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement cites this paper.

UniCSG: Unified High-Fidelity Content-Constrained Style-Driven Generation via Staged Semantic and Frequency Disentanglement Ovis-U1 Technical Report

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:33:41.693409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T05:15:04.561573Z digest=sha256:aa7b27a3fe47131538b4a7a91e7c0ea99f41cea879b645557d4cdf624c603c41

Observation d6f94271-ac7c-42dd-993c-e634b5a8b48d · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ovis-U1 Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:41:18.983968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T04:31:26.325118Z digest=sha256:fde0c6f012ef774c4bc3c5f5baee0726bf4cd966204a8a0a4a3ee88c70051c05

Observation ed7f43f2-3360-479d-bc81-2260ec90bc9f · inbound

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation cites this paper.

Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Ovis-U1 Technical Report

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:43:51.147802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T23:41:25.275207Z digest=sha256:4fe6f984ed1c6112dffcf23e53143087c67d9e681f7cb1a428c90b036aca6743

Observation 3efdaf0f-fbd8-4c7d-9902-b255e18ebbb9 · inbound

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning cites this paper.

UniPath: Adaptive Coordination of Understanding and Generation for Unified Multimodal Reasoning Ovis-U1 Technical Report

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:04.950923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T01:29:23.774196Z digest=sha256:9c51d0553c70e26fe6070426786641a2b590d55618f134fffd69859e433f6a64

Observation aadd1ce5-a9dd-4642-93a2-69c3c47df8ba · inbound

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture cites this paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Ovis-U1 Technical Report

Reference 130

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.564715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:fdf6bbbca1d95d5cd313651a643a49296225239b5c02611a12b5eea283d21345

Observation bfd24f7b-55d3-4bca-9c8e-6d77ce01bf63 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Ovis-U1 Technical Report

Reference 112

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:48:14.989776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-20T11:46:52.658984Z digest=sha256:566e76852496396338f16e0de7c028c1d2e65c129fb8e22a7ecd870084a75b7e

Observation 18da38ba-8842-46f5-a654-b8adc6741fc5 · inbound

Lance: Unified Multimodal Modeling by Multi-Task Synergy cites this paper.

Lance: Unified Multimodal Modeling by Multi-Task Synergy Ovis-U1 Technical Report

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-21T07:59:50.533259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-21T07:56:34.034047Z digest=sha256:15c9dd92e98e7eac888f6d7ae1c94a73ba2c03663da791ec4e39d7fe45f49ad3

Observation fc274802-491c-4c43-a684-3257e5aeeccf · inbound

ProductWebGen: Benchmarking Multimodal Product Webpage Generation cites this paper.

ProductWebGen: Benchmarking Multimodal Product Webpage Generation Ovis-U1 Technical Report

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:06:14.774423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T17:26:44.297446Z digest=sha256:6d2f3880e56606677780aaf868eaa77bc5b6d1d70f148a06e5867afa254e98ba

Observation 02cbe20d-9d72-462f-b43a-dd01b43dd8d6 · inbound

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing cites this paper.

Is This Edit Correct? A Multi-Dimensional Benchmark for Reasoning-Aware Image Editing Ovis-U1 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T20:04:25.886330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:04:25.886330Z digest=sha256:4d552d14ccbfd2c95210f55b09ed128b7a4bdf8ed5719d98913c5b61eeb54e66

Observation 1aa08193-def3-42b2-85cf-1652be0eb18c · inbound

InterleaveThinker: Reinforcing Agentic Interleaved Generation cites this paper.

InterleaveThinker: Reinforcing Agentic Interleaved Generation Ovis-U1 Technical Report

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T15:08:32.893829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T06:42:34.126336Z digest=sha256:e329d5891ac5a0071cc0e72ba645dce5f3e7b540ddaf3ff25848966c2ab665cb

Observation 86578b31-29d0-469b-b3f0-e0c6be404263 · inbound

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? cites this paper.

ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing? Ovis-U1 Technical Report

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:39:17.483974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T21:02:24.792139Z digest=sha256:eaf897a86f1a08ae86009bbc53ee28db9b8c6b4366cf55594d40a393b80e53fd

Observation ee6c667c-5e32-4925-b143-4c580b0bc9b1 · inbound

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget cites this paper.

Boogu-Image-0.1: Boosting Open Agentic Multimodal Generation via Understanding under a Minimal Budget Ovis-U1 Technical Report

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-02T06:14:04.110103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:14:04.110103Z digest=sha256:d46af12b529667bae78ecf08de51b253b7a3d38747121129ffdfd7d162e1a433

Observation 85c6cb5c-2878-4b36-bba8-fc6a7f6c33e4 · inbound

SciForma: Structure-Faithful Generation of Scientific Diagrams cites this paper.

SciForma: Structure-Faithful Generation of Scientific Diagrams Ovis-U1 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T16:12:23.583355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:12:23.583355Z digest=sha256:1caeaf90a9c381ab1062bebf35df89cc26b628963f42b1688dfe2a525b74cdb8

Observation 7feddf75-2d9c-4cbd-9b90-0bcd4e1bda71 · inbound

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing cites this paper.

Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Ovis-U1 Technical Report

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-01T13:39:02.846343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:39:02.846343Z digest=sha256:c6da2fbf568074a0a860e8fd2f803e77b1af47476420f6ccd662f71e7f8c8543

Observation 67dbe261-a3d7-43e7-9ac1-134dc8cd0f8f · inbound

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model cites this paper.

Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model Ovis-U1 Technical Report

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T06:20:13.828990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T06:20:13.828990Z digest=sha256:fd5e683f0aa98ea9dcb6568d8382afe5065cc271f58086807d35c51aec07da30

Observation 808d42c2-22d5-4d08-a6a9-fbfcffebc0da · inbound

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications cites this paper.

Instruction-based Image Editing: A Survey on Data, Models, Evaluation, and Applications Ovis-U1 Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T01:51:50.260499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:51:50.260499Z digest=sha256:d5c6642b80e346ec9efa31649260056d132d61d66823d75d95683567667880b5

Observation cbf59592-eca8-4bfc-9597-a5d824d2a55b · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Ovis-U1 Technical Report

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-06T11:55:28.131719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:55:28.131719Z digest=sha256:6092f180c2e0ce82b7bbed5dbbca12885aca22604ce53932d73404cd6f151b17

Observation d1086410-3241-4a04-985b-2c38ff909417 · inbound

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes cites this paper.

Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Ovis-U1 Technical Report

Reference 123

Resolution
unresolved
no resolver link, observed 2026-08-08T17:08:56.261632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T17:08:56.261632Z digest=sha256:6bf9a09604de3c9a7047fc8377187e6e35542e68ad480791a5d5e31964b60fc9

Observation 23ad3e4c-6cc5-435c-ad12-8a2e9c55ece3 · inbound

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling cites this paper.

UniSpace: Unified Visual Representation and Scalable Multimodal Modeling Ovis-U1 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T04:35:00.046640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:35:00.046640Z digest=sha256:d600b98cbaa3233b1da113c0798aaa989b42e4eb90cce601def525a090d989db

Observation 3f9d7d50-5a78-435e-817c-0e09c284d389 · inbound

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems cites this paper.

A Model-Internal Protocol for Assessing Multimodal Models as Integrated Systems Ovis-U1 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.911595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.911595Z digest=sha256:0dfd981df77f33535ab25f10038e83a977388db29a918c24591aa6936fbd5ce6