Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 100 of 173 outbound references and 15 inbound Pith citation observations for arXiv:2605.12500.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T04:17:41.402174Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T02:04:26.362870Z
100 of 173 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 1f4acabf-572c-4638-aa06-a34a8129976b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flamingo: a visual language model for few-shot learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 104bddc6-2e96-497b-9350-c5dfb4399823 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e3742a16-86be-4c8e-a868-4d3c247d5541 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Imagen 3
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8e998da3-23bf-4739-a366-62998a19074f · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Introducing our multimodal models
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0a484062-6dee-4232-9ee5-fc2c2cb62350 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Improving image generation with better captions
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fd0f4b83-43db-4a66-8d8e-efa064b3709c · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Seedream 4.5
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1595dcbd-9599-46a1-a445-a10c6da4c2a8 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 19a68527-dd97-4535-9fd3-19439d270e91 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0851aadd-5583-42a8-a682-06f9b5b42945 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Has gpt-5 achieved spatial intelligence? an empirical study
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 946f2202-1768-4755-81df-5355968f8f5b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Scaling spatial intelligence with multimodal foundation models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4aba686b-0fa6-4085-a1b0-ab7f8b4af776 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture HunyuanImage 3.0 Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 71d5c1ad-1931-4f39-988a-69d7aa632b8f · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d12a3c72-d5b6-4632-b541-f6ae0ca03266 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture arXiv preprint arXiv:2509.25162 (2025) 4
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation cb6c6fd3-56b9-4952-987d-dc12d54f630a · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2eecc1b5-7e55-4014-b21c-e4c0d30ef38b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BabyVision: Visual Reasoning Beyond Language
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 0ae41b3c-c004-49b2-90eb-e3d964c85286 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37:27056–27087
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f495c277-ff77-46fc-a33a-c0285f78f7df · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture PixelFlow: Pixel-Space Generative Models with Flow
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b6ce9a3f-84d6-475b-bcd1-3ce4aadb8ac0 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e05ac1b0-8a00-447e-9527-580c0425a742 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture A single transformer for scalable vision-language modeling.Transactions on Machine Learning Research
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a4291a7-35e8-4184-93c7-54ea208fe93f · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture arXiv preprint arXiv:2511.18822 (2025)
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9f5be935-48d2-433d-8137-702321f38ce5 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1857c167-9b5c-4916-913a-b1ed054df8d6 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture PaddleOCR 3.0 Technical Report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 15f8e895-b4ee-41ae-abcc-e9c546de505b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emu3.5: Native Multimodal Models are World Learners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9f12aa6a-6822-43f5-bb1d-4b552b9d01e0 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 3 pro image model card, Nov 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4c991b3b-629b-40c6-947a-d66b0b67d4c1 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 2.0 flash
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 659920e2-20a9-40e7-8e91-843b0e633f56 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 2.5 flash & gemini 2.5 flash image model card
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ed3e9c13-99ca-440e-9b0d-b61e561b7be2 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemma 4: Byte for byte, the most capable open models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f0bf9667-6cbf-43ae-9f64-b1a2c94337cc · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emerging Properties in Unified Multimodal Pretraining
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d29c59fa-dbeb-4925-b0fa-1a06aff5717c · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unveiling encoder-free vision-language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3ecc11cd-201a-4d42-ad50-bd6746178fbf · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5b4a9577-7806-441f-ad61-b19bc3a98ca1 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Evev2: Improved baselines for encoder-free vision-language models
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b06e328e-7ad2-44d0-9571-2bbcb51099a7 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Textcrafter: Accurately rendering multiple texts in complex visual scenes
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1a555176-9818-49a9-b386-ae0d5eda1945 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6e84b32d-5411-4541-ac03-61fc41da9ab4 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Scaling rectified flow transformers for high-resolution image synthesis
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 05cedc28-4d04-47ec-b11e-911515f64702 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 04eafcfc-4ee4-451f-866a-27ae18004a14 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Phased dmd: Few-step distribution matching distillation via score matching within subintervals.arXiv preprint arXiv:2510.27684
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a11a964b-2432-49fb-8b43-141d99e97590 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 96d34c76-bfef-4ea5-ad22-f3d1a1c4a295 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Seedream 3.0 Technical Report
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f43ffe58-d574-4613-868b-02f1b8e1feb9 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Making LLaMA SEE and Draw with SEED Tokenizer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ea04040f-fb56-4c54-9a00-b72e7a0b557e · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4c038810-2117-44c8-86ec-2bba18be15cf · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unresolved cited work
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bc1d692f-d031-4224-899e-ba7e78be7c93 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a47bf840-3cc3-467e-9360-54c22cff79a4 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Geneval: An object-focused framework for evaluating text-to-image alignment
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2e82c0ea-db84-4a1c-9282-a6834c7b83ab · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Past-future scheduler for llm serving under sla guarantees
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d7772b3d-7277-40a4-a837-a5e86e59c660 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Imagen 4 model card
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 936dab28-9af6-49cb-8a69-3912342d1ca4 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Nano banana 2: Combining pro capabilities with lightning-fast speed
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9938e232-8458-427a-a576-07bb5f55258e · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 3 Pro Model Card, November 2025
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation c5f5654f-111c-48d0-8e4e-1f7de304f871 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e8470115-ad17-40d5-99a7-4f74c984021f · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b8ca05a5-5a27-4210-b54b-344b35c956ee · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9289e161-1dd6-49ca-b077-6e72100fcac2 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Ai2d-rst
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 910d930d-9505-4e90-aa85-34bfabe303fc · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fb3583f9-1a7b-4a95-a691-e7de0c5e2c8d · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f84d59dc-7e92-4d70-942f-7d3ea52f80a0 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e5be6e8d-3f39-485d-bae3-06b5f27da7cd · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e522aa40-2147-47da-9893-ae08d904c8ad · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture GPT-4o System Card
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6cb9f8bd-c4b9-4930-b335-80cc611e1ac8 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Kimi K2: Open Agentic Intelligence
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6ad436d7-7d8b-4bf4-9533-f7168ad9429d · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Auto-Encoding Variational Bayes
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8b786f63-54fe-4772-ac95-5d593950379b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Kolors 2.0
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 98617361-80fd-4e51-a325-9850284f742c · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flux
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 36c687be-eb70-4217-9214-1d4cde96c3b3 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture FLUX.2: Frontier Visual Intelligence
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 315466c3-f697-4cb2-ac1f-b10cafd405c9 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1d606c1f-2b99-4180-9835-63c350a1e62f · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 32e3dd20-cf57-4c66-8626-991e50ba65f3 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f2c28161-d97e-4b6e-b4d3-d2b20cbaa561 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.ArXiv, abs/2505.21500
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 30a5e343-3429-4d2f-bce0-7645e513a7e0 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Onecat: Decoder-only auto-regressive model for unified understanding and generation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1b8e3371-fe0d-4001-9ceb-217e5649f6e9 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 88daa041-5906-4682-9b74-4d897a824f49 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tir-bench: A comprehensive benchmark for agentic thinking-with-images reasoning
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8a364973-830e-4f19-974a-a05d4a268d0e · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Back to Basics: Let Denoising Generative Models Denoise
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7ca87196-2182-475a-9d2d-bdf09c9d34f7 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Breen: bridge data-efficient encoder-free multimodal learning with learnable queries
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ebd8ae8d-a66f-45c3-bcc2-85f99d1c488d · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Bizgeneval: A systematic benchmark for commercial visual content generation
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7de6334b-e280-48a8-8b50-ae5cdd89c236 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8490c6c8-47a4-49d3-baea-426d2faccba8 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 549d4ce2-198c-4882-9690-89277b67842b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e902012f-d892-42c9-95a1-1a5d2f5a713a · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3a157326-a4d9-4ee4-a2bd-2f5483acdd24 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 342bab0a-83bf-49d1-8305-6ecda8da66b5 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Microsoft coco: Common objects in context
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 57e87d1c-9fd2-41f0-bc2f-3bfdd81001db · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 252c684e-3c08-469c-8dce-64d8fe1883f8 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Visual instruction tuning
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 8bd0a807-3759-4777-a36f-8df649eea968 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flow-GRPO: Training Flow Matching Models via Online RL
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1dfcc8d3-cc39-4dd3-951c-269755acd279 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Step1X-Edit: A Practical Framework for General Image Editing
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 896535d9-5dd5-463c-a14b-e108a59cfa43 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, pages 216–233
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9967e967-951c-48c3-932c-ed9d649f6ad0 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Ocrbench: on the hidden mystery of ocr in large multimodal models
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 50b46e35-daa3-462c-8d04-c1f577a878e2 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna: Taming unified visual representations for native unified multimodal models
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7e87a0de-826d-44e1-b9fd-4df531930344 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1d2cdf16-c3a8-4b6a-a9d3-58525c4974ee · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Stable diffusion 3: Research paper
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 41bbb502-84f4-4c01-95dd-807771edde84 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 85ce1315-55db-4859-9f18-b28e63813223 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation b7113530-4fdf-4d9f-b603-0d3e3ad84bdc · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 99796e2b-f2c4-491c-93be-267ba69516dc · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2af374c7-54e3-4a15-81a0-0d793229f689 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 3dsrbench: A comprehensive 3d spatial reasoning benchmark
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5aba762f-585f-4f07-8ff7-2eb50c825c9d · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 05195b58-2b82-49fc-92f2-6e959aa97f48 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Hpsv3: Towards wide-spectrum human preference score
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 90d7109a-3a0c-450c-afbe-acb5f96e6381 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Infographicvqa
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 30fe1491-7edd-4187-a310-0a6563afafb4 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Midjourney v7
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 08fc4245-b99a-4b20-8e68-1840c915859a · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Image-01
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 4fcd9f5c-e310-4b12-a4e7-f9d7fed13764 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture LightLLM: A python-based llm inference and serving framework
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 59ceda2b-ced0-4a3e-990c-277750e74d3b · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture LightX2V: A lightweight video and image generation inference framework
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 84c3964e-28d5-42c8-b513-7e1841cf4f37 · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ea26a3db-4287-4eac-b9ea-8b6e0c69d1fc · outbound
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gpt-image-1
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 7397f622-e740-4968-ad16-2d500d2353f7 · inbound
Toward Native Multimodal Modeling: A Roadmap SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 78146988-1f12-4e74-b914-132dd2518d34 · inbound
How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ac6bc258-b7d3-4f08-b0bf-94cf4f60932c · inbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation eec1c095-5b55-49ab-8e4e-e111cafe4944 · inbound
Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e62b9711-82d3-4ed2-ae0c-cb8bbada6876 · inbound
Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2875229e-e963-4ef7-ac89-44815a95dc0a · inbound
WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 482ae9cd-c5df-4a8d-a343-ee798656f161 · inbound
ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3b0384e4-509b-42e6-ad03-e48e446b8c9b · inbound
GEAR: Guided End-to-End AutoRegression for Image Synthesis SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2c37936e-c6c7-41b2-bcd1-c810c43c906f · inbound
DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation fc9d60c5-a1f8-4468-a9f1-a5ba17bf8b9e · inbound
Vision as Unified Multimodal Generation SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9cace125-3dd5-41a2-88e9-ce4217fe4fae · inbound
SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4579cabf-9355-4505-80c7-f018bb3855eb · inbound
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb9eaa3-f7b0-4224-8e16-8cf6f3a294b7 · inbound
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37a337e9-960c-47a3-aee4-a1388418953c · inbound
STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2401b0-516e-4693-a4ea-365964f7b622 · inbound
Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.