Pith. sign in

Paper Citation Record · LEDGER

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

As of 4 August 2026, this Paper Citation Record lists 100 of 173 outbound references and 15 inbound Pith citation observations for arXiv:2605.12500.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12500 v1

Coverage vector

measured 100 of 173 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T05:12:37.339084Z

measured 115 of 115 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T04:17:41.402174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-08T02:04:26.362870Z

Reference resolution

100 of 173 outbound references displayed

  • verified exact51
  • verified fuzzy44
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch4

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1f4acabf-572c-4638-aa06-a34a8129976b · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.506554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:d3a8d49c2e8cb9ec41601a29ef25f0c959600c3af52d359015b8c7fbb16a56b7

Observation 104bddc6-2e96-497b-9350-c5dfb4399823 · outbound

This paper cites Qwen3-VL Technical Report.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.807782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:f768192f0e4e18f5b5268bfcfad4298c27b77bc4bd628ebc31212bcac4192a97

Observation e3742a16-86be-4c8e-a868-4d3c247d5541 · outbound

This paper cites Imagen 3.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Imagen 3

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.480326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:cd5f2c84015edf3285bccf91bc4da367bbf51bddb40439369aef77c0c78da719

Observation 8e998da3-23bf-4739-a366-62998a19074f · outbound

This paper cites Introducing our multimodal models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Introducing our multimodal models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.510060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:d2db46e94d64a05c9ffda56d1792475b6a55205292b95fba6aa2322ac5cc4575

Observation 0a484062-6dee-4232-9ee5-fc2c2cb62350 · outbound

This paper cites Improving image generation with better captions.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Improving image generation with better captions

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.543956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:cdee5d726908ce80e77d6cc8bafcf67e2c38d9358469af6d6804bb5e60b45391

Observation fd0f4b83-43db-4a66-8d8e-efa064b3709c · outbound

This paper cites Seedream 4.5.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Seedream 4.5

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.470808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:e6f4ad10c293f95cff33db6a01249d6f1cac633e97376fc0327ad5881f42c7e9

Observation 1595dcbd-9599-46a1-a445-a10c6da4c2a8 · outbound

This paper cites Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.556563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:b01559afd311de3f19278b90090ff709630c779ec2c6ab1a03b9d7d3b74c0843

Observation 19a68527-dd97-4535-9fd3-19439d270e91 · outbound

This paper cites HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture HiDream-I1: A High-Efficient Image Generative Foundation Model with Sparse Diffusion Transformer

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T17:13:39.137651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:3f5151a3b0560a2e08c4e39e958d3b9419acd05b29871ab201bc0b50c9ca6f76

Observation 0851aadd-5583-42a8-a682-06f9b5b42945 · outbound

This paper cites Has gpt-5 achieved spatial intelligence? an empirical study.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Has gpt-5 achieved spatial intelligence? an empirical study

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.785885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:0fe6191db1aa8c6160b27f4b41595cc180e387709d5acdda5249ca338e8fa531

Observation 946f2202-1768-4755-81df-5355968f8f5b · outbound

This paper cites Scaling spatial intelligence with multimodal foundation models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Scaling spatial intelligence with multimodal foundation models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.472623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:d53aa865b40f88b45dcf1369c78e8f3f343b998c033d7a65d885af2bc3bdbac7

Observation 4aba686b-0fa6-4085-a1b0-ab7f8b4af776 · outbound

This paper cites HunyuanImage 3.0 Technical Report.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture HunyuanImage 3.0 Technical Report

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:32.945705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:63ed9edcfb8840cd666c8446234f5428ce2b7b8f079475dbec49696a3981cfd7

Observation 71d5c1ad-1931-4f39-988a-69d7aa632b8f · outbound

This paper cites OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OneIG-Bench: Omni-dimensional Nuanced Evaluation for Image Generation

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.581960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:2d4156111326c60d38f35390710ce183a52a45fabd3d996439c53f18314ed470

Observation d12a3c72-d5b6-4632-b541-f6ae0ca03266 · outbound

This paper cites arXiv preprint arXiv:2509.25162 (2025) 4.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture arXiv preprint arXiv:2509.25162 (2025) 4

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.570432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:b949c7334d15aa3026ebffb7ebc67d344da40a07702198d80bc90c821f9743f6

Observation cb6c6fd3-56b9-4952-987d-dc12d54f630a · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.725291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:5cb41a4590137a3962b2594b7140d1eb4656440f881c413eb703751c5f5018ed

Observation 2eecc1b5-7e55-4014-b21c-e4c0d30ef38b · outbound

This paper cites BabyVision: Visual Reasoning Beyond Language.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture BabyVision: Visual Reasoning Beyond Language

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-08T02:18:40.971166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:abeb24d0ad71c1975ef051b19a677933e56e99929e478c7ea13ec1718ed6bc04

Observation 0ae41b3c-c004-49b2-90eb-e3d964c85286 · outbound

This paper cites Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37:27056–27087.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Are we on the right way for evaluating large vision-language models? Advances in Neural Information Processing Systems, 37:27056–27087

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.511912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:05ed25265533b9a6ccd9ecb81a588bc26c287a96c5e6dc017a9b507b972c669b

Observation f495c277-ff77-46fc-a33a-c0285f78f7df · outbound

This paper cites PixelFlow: Pixel-Space Generative Models with Flow.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture PixelFlow: Pixel-Space Generative Models with Flow

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.498517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:b2f051190f134fbda73585c983c0872bed3660c7739ba7fb607489c5d1ee2eb7

Observation b6ce9a3f-84d6-475b-bcd1-3ce4aadb8ac0 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.728349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:5e0e085020eaec563dfca0766f8f02d69829e3f61d1f1915d2ff5c7f5d4d9504

Observation e05ac1b0-8a00-447e-9527-580c0425a742 · outbound

This paper cites A single transformer for scalable vision-language modeling.Transactions on Machine Learning Research.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture A single transformer for scalable vision-language modeling.Transactions on Machine Learning Research

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.474488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:52620915bd8bfeaa4b15a8537f3db71e88e7e5af397c7b1d7e60d8be9623fc59

Observation 1a4291a7-35e8-4184-93c7-54ea208fe93f · outbound

This paper cites arXiv preprint arXiv:2511.18822 (2025).

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture arXiv preprint arXiv:2511.18822 (2025)

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.701506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:94c976294c8b5b47ad5af2ca1231849c3f0d2610d38c62366d6db6c8e052622b

Observation 9f5be935-48d2-433d-8137-702321f38ce5 · outbound

This paper cites ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ANOLE: An Open, Autoregressive, Native Large Multimodal Models for Interleaved Image-Text Generation

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.811647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:db21e195f496489bf901a0b4f65736a4beab120c9c6db7ae175d43f673e91b04

Observation 1857c167-9b5c-4916-913a-b1ed054df8d6 · outbound

This paper cites PaddleOCR 3.0 Technical Report.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture PaddleOCR 3.0 Technical Report

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-14T23:24:28.568269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:2170f0bc185655e17916d7e2b60194a17973b091ecc9457794349e5de2e47b96

Observation 15f8e895-b4ee-41ae-abcc-e9c546de505b · outbound

This paper cites Emu3.5: Native Multimodal Models are World Learners.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emu3.5: Native Multimodal Models are World Learners

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:12:13.931083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:4bcf841d46f24370693900131a62c332ad74cdc2bc3b1aa8a199882b458ca52d

Observation 9f12aa6a-6822-43f5-bb1d-4b552b9d01e0 · outbound

This paper cites Gemini 3 pro image model card, Nov 2025.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 3 pro image model card, Nov 2025

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.479601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:389b9d476b938862b8a03ed4bbe4b0629795a75644cad3485c585c9539dec8dd

Observation 4c991b3b-629b-40c6-947a-d66b0b67d4c1 · outbound

This paper cites Gemini 2.0 flash.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 2.0 flash

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.477915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:9fb3974286f217d099e659000afb936801b0b1fce11ac6b147fa3b49fd935ed5

Observation 659920e2-20a9-40e7-8e91-843b0e633f56 · outbound

This paper cites Gemini 2.5 flash & gemini 2.5 flash image model card.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 2.5 flash & gemini 2.5 flash image model card

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.535110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:c77d004d5564d66d3e2ddeae632a2c5c9ed26dcb1a7fbc5766b46c9b16870f9a

Observation ed3e9c13-99ca-440e-9b0d-b61e561b7be2 · outbound

This paper cites Gemma 4: Byte for byte, the most capable open models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemma 4: Byte for byte, the most capable open models

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.469120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ec4265eb552def2b84a32265179cd0f1d4edb5008f43a4a05f3a4193d89ca328

Observation f0bf9667-6cbf-43ae-9f64-b1a2c94337cc · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Emerging Properties in Unified Multimodal Pretraining

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.694161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:5f8afd4e2361d5fccc8b571c8f84add738cb404905e3e116f9847d110969dac1

Observation d29c59fa-dbeb-4925-b0fa-1a06aff5717c · outbound

This paper cites Unveiling encoder-free vision-language models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unveiling encoder-free vision-language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.461578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:25d0912a10909079d9c1045714044dc9143de703213cbc68251c51979c0a4246

Observation 3ecc11cd-201a-4d42-ad50-bd6746178fbf · outbound

This paper cites From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture From pixels to words–towards native vision-language primitives at scale.arXiv preprint arXiv:2510.14979, 2025a

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.545969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:2a3680164af03a5cb4fd02a99b645d8f04ead3f40c16beef51d94cf2c6a58b63

Observation 5b4a9577-7806-441f-ad61-b19bc3a98ca1 · outbound

This paper cites Evev2: Improved baselines for encoder-free vision-language models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Evev2: Improved baselines for encoder-free vision-language models

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.465470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:9e6aea3e0ed4cca7f3aa8861aa6d42dc4dd579673ea1291fd6eebfdda03042e1

Observation b06e328e-7ad2-44d0-9571-2bbcb51099a7 · outbound

This paper cites Textcrafter: Accurately rendering multiple texts in complex visual scenes.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Textcrafter: Accurately rendering multiple texts in complex visual scenes

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.690826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:f0fe1e26e8ca16d60e18c75ec6c4239821ee73000d5db6e24a38829640833c05

Observation 1a555176-9818-49a9-b386-ae0d5eda1945 · outbound

This paper cites SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T00:47:04.366365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:98d3cd3d302dda6a87e97cf36fbd6465b1e58bb10a9c699b3dc15e46cf59e7e8

Observation 6e84b32d-5411-4541-ac03-61fc41da9ab4 · outbound

This paper cites Scaling rectified flow transformers for high-resolution image synthesis.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Scaling rectified flow transformers for high-resolution image synthesis

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.529324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a158bc1095e5f9d0d6928c8014986ba4f55f444a343945dd6e3a1201d37b6f36

Observation 05cedc28-4d04-47ec-b11e-911515f64702 · outbound

This paper cites The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture The prism hypothesis: Harmonizing semantic and pixel representations via unified autoencoding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.745699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a2a9425b771d1c64ab38c37486bde705f38fa24bba9b9a569082a93444070a35

Observation 04eafcfc-4ee4-451f-866a-27ae18004a14 · outbound

This paper cites Phased dmd: Few-step distribution matching distillation via score matching within subintervals.arXiv preprint arXiv:2510.27684.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Phased dmd: Few-step distribution matching distillation via score matching within subintervals.arXiv preprint arXiv:2510.27684

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.712253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:60cac33d6cb83a37099e2bee06ac2477de8e4a3309230945b054cc14be390205

Observation a11a964b-2432-49fb-8b43-141d99e97590 · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:33:27.157097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:fa0e23b845df78d3ecf5fcf180454b3dd3a121337ca4d71ebadda8fc2070456a

Observation 96d34c76-bfef-4ea5-ad22-f3d1a1c4a295 · outbound

This paper cites Seedream 3.0 Technical Report.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Seedream 3.0 Technical Report

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:55:38.865721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:0c6f43a1100124a8e7ff883e2917e198121689d0df8640e3e7133f809bdefe2a

Observation f43ffe58-d574-4613-868b-02f1b8e1feb9 · outbound

This paper cites Making LLaMA SEE and Draw with SEED Tokenizer.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Making LLaMA SEE and Draw with SEED Tokenizer

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.821835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:7a309c049b68a4e12d785b2ae3622af62a4f874bb1b60f42760aabb9e5a13c64

Observation ea04040f-fb56-4c54-9a00-b72e7a0b557e · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:48:36.519178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:bcb01eb03621a2f7ccd5994ac066a791aa0ba6a946651842812340f698688580

Observation 4c038810-2117-44c8-86ec-2bba18be15cf · outbound

This paper cites an unresolved cited work.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-05-13T10:52:39.499323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:478ca8255b167a2591e72693becab2126a225e4fdf9716dbe26368f7686f9eca

Observation bc1d692f-d031-4224-899e-ba7e78be7c93 · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.603130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:62ab11702f51816b591f28d6ee90531a586d3b3633c7e91fd74c8aa3542bad18

Observation a47bf840-3cc3-467e-9360-54c22cff79a4 · outbound

This paper cites Geneval: An object-focused framework for evaluating text-to-image alignment.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Geneval: An object-focused framework for evaluating text-to-image alignment

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.444983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:492dfe5fba0abb2bdb3a676747b07836ec6c62def552bdc46bc0b6227b570037

Observation 2e82c0ea-db84-4a1c-9282-a6834c7b83ab · outbound

This paper cites Past-future scheduler for llm serving under sla guarantees.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Past-future scheduler for llm serving under sla guarantees

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.443118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:80544e165ab4a720b069266e1644405d0d8bde3f5becc235ee4cee17213bbb41

Observation d7772b3d-7277-40a4-a837-a5e86e59c660 · outbound

This paper cites Imagen 4 model card.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Imagen 4 model card

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.446664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:6e70579c611b254606eb128a7e9d9a72c268d74051587cab7e88218138b1d550

Observation 936dab28-9af6-49cb-8a69-3912342d1ca4 · outbound

This paper cites Nano banana 2: Combining pro capabilities with lightning-fast speed.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Nano banana 2: Combining pro capabilities with lightning-fast speed

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.450217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:55cf689cc688ae5fca6460618e0cebcf942b791997ff4bcbddbb58480bbcf4de

Observation 9938e232-8458-427a-a576-07bb5f55258e · outbound

This paper cites Gemini 3 Pro Model Card, November 2025.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gemini 3 Pro Model Card, November 2025

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.459739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ffbc8c7ccd9bb26a09c52cdbd68bdb452bebb2a6a0965cc0a7939d0319d4305c

Observation c5f5654f-111c-48d0-8e4e-1f7de304f871 · outbound

This paper cites Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Thinkmorph: Emergent properties in multimodal interleaved chain-of-thought reasoning

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.483141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:0c8dee16762cae259c67f35f2eb9941976835784d75303cfca779244c09a3f25

Observation e8470115-ad17-40d5-99a7-4f74c984021f · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.458016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:3351eda9643d3a60aa0a7369e260ca1206bb8b79b07a18636ea81ccd64d50771

Observation b8ca05a5-5a27-4210-b54b-344b35c956ee · outbound

This paper cites Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Infinity: Scaling bitwise autoregressive modeling for high-resolution image synthesis

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.435630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a0c2adb33b13f65be999cf6eafc7e1e97ae166b2a9c023e562783c9ede1e3fd7

Observation 9289e161-1dd6-49ca-b077-6e72100fcac2 · outbound

This paper cites Ai2d-rst.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Ai2d-rst

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.437628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a5d3dd263661c147157943e730aaf24e3b04d3503073acbd742a4928feee2275

Observation 910d930d-9505-4e90-aa85-34bfabe303fc · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.674337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:9ea0a03695f1c1edf25bb79c3d248d87820546f5dea1a6ed60c24046fc0d0961

Observation fb3583f9-1a7b-4a95-a691-e7de0c5e2c8d · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.760218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:098f0564b69e193ba014c8c92900c7926f005b25edb3273564bef950cadfa269

Observation f84d59dc-7e92-4d70-942f-7d3ea52f80a0 · outbound

This paper cites ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture ILLUME+: Illuminating Unified MLLM with Dual Visual Tokenization and Diffusion Refinement

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.739464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:db90f9bc39464a8232beb3df5b1e424a3bd04f2fcc9744d8947fd1e9f92ca11b

Observation e5be6e8d-3f39-485d-bae3-06b5f27da7cd · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.439459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:fa96606d50e7c31685aa8f86500f2ef79cc3937c081ce468f521c8d1a6321490

Observation e522aa40-2147-47da-9893-ae08d904c8ad · outbound

This paper cites GPT-4o System Card.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture GPT-4o System Card

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.756274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:e9f329c7fb182d9a8c8ce6cad07db717bdf97ff945d66f5598909bc6198abf11

Observation 6cb9f8bd-c4b9-4930-b335-80cc611e1ac8 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Kimi K2: Open Agentic Intelligence

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.735424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:0e8699e65fad3aa32cf3da3fc8f7e0efffe380f1403a7afec20a6d37c0e6335d

Observation 6ad436d7-7d8b-4bf4-9533-f7168ad9429d · outbound

This paper cites Auto-Encoding Variational Bayes.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Auto-Encoding Variational Bayes

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.532228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:037058618b99d585ca703bc4b3a0bdcfdbd25823d298796c618e375f37224f6c

Observation 8b786f63-54fe-4772-ac95-5d593950379b · outbound

This paper cites Kolors 2.0.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Kolors 2.0

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.502836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:04b272fb6e9df47fc6aa8802197147147fec16486b6aa64354c61a4a6c759ef0

Observation 98617361-80fd-4e51-a325-9850284f742c · outbound

This paper cites Flux.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flux

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.538644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:357ba4aca8eeae215ad9c6544801092d8560842c46ab887e6184db2204c39695

Observation 36c687be-eb70-4217-9214-1d4cde96c3b3 · outbound

This paper cites FLUX.2: Frontier Visual Intelligence.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture FLUX.2: Frontier Visual Intelligence

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.545613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:493308098bcec6d41fa6d9e3a96ab2b31bddb1a8c72cb7aa3acfbca685bbd697

Observation 315466c3-f697-4cb2-ac1f-b10cafd405c9 · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.526215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:6e8fdd1ae73f0667b12ee586592b32f9e193c14e6a54a04207a9be6750886cd1

Observation 1d606c1f-2b99-4180-9835-63c350a1e62f · outbound

This paper cites The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.454021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:25e834f87caa6d1f2b818488e6493f3e9128e80f11039bb54c2f7366d94a0cb2

Observation 32e3dd20-cf57-4c66-8626-991e50ba65f3 · outbound

This paper cites Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Repa-e: Unlocking vae for end-to-end tuning of latent diffusion transformers

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.456131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:eca373a21a951354f0f0a739c55c541a9f703ba106748aab587a6502ead00fce

Observation f2c28161-d97e-4b6e-b4d3-d2b20cbaa561 · outbound

This paper cites Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.ArXiv, abs/2505.21500.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Viewspatial-bench: Evaluating multi-perspective spatial localization in vision-language models.ArXiv, abs/2505.21500

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.537844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:26bcd3929a65203081919da0b1da630fc09f051185428c99b2627500d1d75fc3

Observation 30a5e343-3429-4d2f-bce0-7645e513a7e0 · outbound

This paper cites Onecat: Decoder-only auto-regressive model for unified understanding and generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Onecat: Decoder-only auto-regressive model for unified understanding and generation

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.486368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:1b672e0214d8c10d3e40a40a1f95a8abed291f3fdedcc2ec6912e02bf7b1910f

Observation 1b8e3371-fe0d-4001-9ceb-217e5649f6e9 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.433697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:96f06af99c9d64e4f514ca7493818a3c3982f85183107f3743ba2fbc051975d4

Observation 88daa041-5906-4682-9b74-4d897a824f49 · outbound

This paper cites Tir-bench: A comprehensive benchmark for agentic thinking-with-images reasoning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tir-bench: A comprehensive benchmark for agentic thinking-with-images reasoning

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.732293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a7b3ef2d6fe767044fb1e737d071b90b85e1b9385cdcdece32d7195b7995600b

Observation 8a364973-830e-4f19-974a-a05d4a268d0e · outbound

This paper cites Back to Basics: Let Denoising Generative Models Denoise.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Back to Basics: Let Denoising Generative Models Denoise

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.800887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:d91342d067c51bf98a5f8baf521d98b99d7419bdc8890e1f070c00a834e32485

Observation 7ca87196-2182-475a-9d2d-bdf09c9d34f7 · outbound

This paper cites Breen: bridge data-efficient encoder-free multimodal learning with learnable queries.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Breen: bridge data-efficient encoder-free multimodal learning with learnable queries

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.441282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:add0aa5767ff53ff1acad14224220919ad03a2ea77e8726bd25d30fb83d6493d

Observation ebd8ae8d-a66f-45c3-bcc2-85f99d1c488d · outbound

This paper cites Bizgeneval: A systematic benchmark for commercial visual content generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Bizgeneval: A systematic benchmark for commercial visual content generation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.492515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:5731df8c4e4aa617e03abca7ae5ee1cfd997c6d24f191fa66de59c9e4aea5d70

Observation 7de6334b-e280-48a8-8b50-ae5cdd89c236 · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:01:19.940440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:50d005383c737b74f9729a44e18110097fc8ad0eb9b104743ba7508a93935656

Observation 8490c6c8-47a4-49d3-baea-426d2faccba8 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-18T02:48:45.158778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:1a5abbab72a1a244d47a4137a6f981a19746a3c4697b6599ca67dfb580fb72dd

Observation 549d4ce2-198c-4882-9690-89277b67842b · outbound

This paper cites Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mogao: An Omni Foundation Model for Interleaved Multi-Modal Generation

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:24:05.047503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:dafc23a9a16df4b661cfb0cc8cdc450f4f058c3de446082582038953cc242d2a

Observation e902012f-d892-42c9-95a1-1a5d2f5a713a · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.517445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:38941c712b742b6e7cc0aceba016cf212343333c426a7ba7bf0ce3b4c856f995

Observation 3a157326-a4d9-4ee4-a2bd-2f5483acdd24 · outbound

This paper cites TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture TokLIP: Marry Visual Tokens to CLIP for Multimodal Comprehension and Generation

Reference 76

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T05:17:18.721952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:22bb38069038e9c6d135ad641e0f5a6c483621c59c0f6d311df386845373e71c

Observation 342bab0a-83bf-49d1-8305-6ecda8da66b5 · outbound

This paper cites Microsoft coco: Common objects in context.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Microsoft coco: Common objects in context

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.463562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:edf595e7e46cebfbfadfa3361bbcfbc38ca763e538462f693931ed8bd7868338

Observation 57e87d1c-9fd2-41f0-bc2f-3bfdd81001db · outbound

This paper cites MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture MoMa: Efficient Early-Fusion Pre-training with Mixture of Modality-Aware Experts

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.708488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:e0a672d36568cbb45584be718f09f7026b65f402c035bb348733a79096cd434e

Observation 252c684e-3c08-469c-8dce-64d8fe1883f8 · outbound

This paper cites Visual instruction tuning.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Visual instruction tuning

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.429665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:5323a0a45ef020f472348936205cf399fff5534d750bdf1c42d0cd4ae15f4941

Observation 8bd0a807-3759-4777-a36f-8df649eea968 · outbound

This paper cites Flow-GRPO: Training Flow Matching Models via Online RL.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Flow-GRPO: Training Flow Matching Models via Online RL

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.804504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ddd6a1e77aabf867858fed65f5c19d0cf087586518d7012c5173e9dd3a0c0d9a

Observation 1dfcc8d3-cc39-4dd3-951c-269755acd279 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Step1X-Edit: A Practical Framework for General Image Editing

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.643765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:10f9263a3797e589ee406dbf90ca597f81da7113346393452c09a2124a2599a3

Observation 896535d9-5dd5-463c-a14b-e108a59cfa43 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, pages 216–233.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mmbench: Is your multi-modal model an all-around player? In Proceedings of the European Conference on Computer Vision, pages 216–233

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.421664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:30a125946491c9dd3a23a100ec2d0e523b4aef0d3cfec8aae78cda1ca2a66def

Observation 9967e967-951c-48c3-932c-ed9d649f6ad0 · outbound

This paper cites Ocrbench: on the hidden mystery of ocr in large multimodal models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Ocrbench: on the hidden mystery of ocr in large multimodal models

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.425818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:8424b438c2c33c54d5cb6761ed626bbf83eab5c0b86261ac16dafae82e0f17ad

Observation 50b46e35-daa3-462c-8d04-c1f577a878e2 · outbound

This paper cites Tuna: Taming unified visual representations for native unified multimodal models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna: Taming unified visual representations for native unified multimodal models

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.567494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:af2f00883bb0ff4e7cbebd9f1ecbce93648c6df18d5529f1dda3f473120855da

Observation 7e87a0de-826d-44e1-b9fd-4df531930344 · outbound

This paper cites Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation

Reference 85

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.656071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:29c225affe13b73fc72b01d77f0c0be3b440645edd8f99c7ad6ea06520376170

Observation 1d2cdf16-c3a8-4b6a-a9d3-58525c4974ee · outbound

This paper cites Stable diffusion 3: Research paper.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Stable diffusion 3: Research paper

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.431536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:ec128660ca252770b7baf1ca4cbb4cf4852cb67b9eb69d61b82c788723dbe144

Observation 41bbb502-84f4-4c01-95dd-807771edde84 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 87

Resolution
verified exact
local_arxiv, observed 2026-05-13T05:17:18.764085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:840241f8d9061f19079a7a4fc0ed37bc3d019d43491b7052923aaaddbe5b661f

Observation 85ce1315-55db-4859-9f18-b28e63813223 · outbound

This paper cites Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.752838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:d21d616da0b38641a520a572f598e6f0f8a6385b107639847689555a866694e3

Observation b7113530-4fdf-4d9f-b603-0d3e3ad84bdc · outbound

This paper cites Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Mono-internvl: Pushing the boundaries of monolithic multimodal large language models with endogenous visual pre-training

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.417939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:7068af3bdc760a484ccdabee518b68cee4b80045989671be02b5555fe88ec8ea

Observation 99796e2b-f2c4-491c-93be-267ba69516dc · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.678621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:c4eaaf49d40f6aa3ba203946ae62dbd385b1d5d9295571fd2fd463425e4ccbc9

Observation 2af374c7-54e3-4a15-81a0-0d793229f689 · outbound

This paper cites 3dsrbench: A comprehensive 3d spatial reasoning benchmark.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture 3dsrbench: A comprehensive 3d spatial reasoning benchmark

Reference 91

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:17:18.609205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:127af5be533eb50c9462273db173da045ba473217b9f277cb40810db1e447955

Observation 5aba762f-585f-4f07-8ff7-2eb50c825c9d · outbound

This paper cites Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Janusflow: Harmonizing autoregression and rectified flow for unified multimodal understanding and generation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.419733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:46e5ac9d892bf4bbbecde776a2d45edfbb896effa6ce1bb25a67b6fed3925392

Observation 05195b58-2b82-49fc-92f2-6e959aa97f48 · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Hpsv3: Towards wide-spectrum human preference score

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.427730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:bdb8645017e27f5ae14ddb328ca9da0817ce6d026b53e84ca20e4a2c8b2c71d2

Observation 90d7109a-3a0c-450c-afbe-acb5f96e6381 · outbound

This paper cites Infographicvqa.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Infographicvqa

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.467321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:80fc34617d82129d1c9afe964e7de4a9efa9ae00f3f35aebfbab40eb75587710

Observation 30fe1491-7edd-4187-a310-0a6563afafb4 · outbound

This paper cites Midjourney v7.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Midjourney v7

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.540340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:052612cd5d00acec97a9a62eefc65697b7e13a542c03af2e5d68ddae8c55a38a

Observation 08fc4245-b99a-4b20-8e68-1840c915859a · outbound

This paper cites Image-01.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Image-01

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.523988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:7139ad868a5324dca825ee2e06e9625d2aa5f10f378decdcf043128df938dcbd

Observation 4fcd9f5c-e310-4b12-a4e7-f9d7fed13764 · outbound

This paper cites LightLLM: A python-based llm inference and serving framework.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture LightLLM: A python-based llm inference and serving framework

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.508356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:559bd86a234c500f57e89b6057577a80264b606782c0b91c06c90e2d187b4704

Observation 59ceda2b-ced0-4a3e-990c-277750e74d3b · outbound

This paper cites LightX2V: A lightweight video and image generation inference framework.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture LightX2V: A lightweight video and image generation inference framework

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.476260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:a47d96daa11e3bb34330bc5092e604ab0bdcefc62c2a01af10466f6f0082fdde

Observation 84c3964e-28d5-42c8-b513-7e1841cf4f37 · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.819650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:1d5f9a562a6217f3e79349eb7568fc02666798f1a9a3307a95b58e75d55c04be

Observation ea26a3db-4287-4eac-b9ea-8b6e0c69d1fc · outbound

This paper cites Gpt-image-1.

SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Gpt-image-1

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-05-13T10:52:39.515485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-13T05:12:37.339084Z digest=sha256:583d27d894a06cac1287318978c309de7ff5450bc90ead5f3fb33878a4fd337e

Pith citing papers

Observation 7397f622-e740-4968-ad16-2d500d2353f7 · inbound

Toward Native Multimodal Modeling: A Roadmap cites this paper.

Toward Native Multimodal Modeling: A Roadmap SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:04:01.929840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-29T22:58:38.610609Z digest=sha256:ba78d10dac9a68414c7f26fe426c5cc5223ea94bc98b94332e343ae006d38cd0

Observation 78146988-1f12-4e74-b914-132dd2518d34 · inbound

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning cites this paper.

How and What to Imagine? Visual Thinking in Unified Multimodal Models for Cross-View Spatial Reasoning SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-06-29T18:23:50.686391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-06-29T18:19:40.323661Z digest=sha256:7d4095c3750b882e71609cda4226110bb3d3047c01df8f6548b8ee370ab751e3

Observation ac6bc258-b7d3-4f08-b0bf-94cf4f60932c · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-01T19:16:00.959654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-28T22:54:10.460872Z digest=sha256:f7c1df1f46b668993eea3d0f4f2b3545d4c75a9a1e0b8e27ba922c1095c108ac

Observation eec1c095-5b55-49ab-8e4e-e111cafe4944 · inbound

Representation Forcing for Bottleneck-Free Unified Multimodal Models cites this paper.

Representation Forcing for Bottleneck-Free Unified Multimodal Models SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T15:31:57.426559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T15:31:57.426559Z digest=sha256:c00bd90f314eed6d2b9cce0edf33193392b962717193ebdb17663b70d9330676

Observation e62b9711-82d3-4ed2-ae0c-cb8bbada6876 · inbound

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion cites this paper.

Show the Signal, Hide the Noise: Spectral Forcing for Pixel-Space Diffusion SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-03T17:18:43.615339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-27T04:21:43.851963Z digest=sha256:e217ff782dabebc1680672af1e3c40a95b70b8466528f68baa47cc858f78e4b1

Observation 2875229e-e963-4ef7-ac89-44815a95dc0a · inbound

WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization cites this paper.

WeGenBench: A Multidimensional Diagnostic Benchmark towards Text-to-Image Model Optimization SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:39:29.350086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-06-26T17:54:09.656061Z digest=sha256:80343f0908ada41ae547df019a7525a6fac8b01a1eeaad80fa432f1c89c08ab7

Observation 482ae9cd-c5df-4a8d-a343-ee798656f161 · inbound

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs cites this paper.

ERA: Entropy-Guided Visual Token Pruning with Rectified Attention for Efficient MLLMs SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:15:44.328438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T05:44:12.758733Z digest=sha256:5ddaf8b1c181890d87841a17e3c15078fcc1f31ffa3a45541ad49e494d0c5f30

Observation 3b0384e4-509b-42e6-ad03-e48e446b8c9b · inbound

GEAR: Guided End-to-End AutoRegression for Image Synthesis cites this paper.

GEAR: Guided End-to-End AutoRegression for Image Synthesis SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-01T10:35:42.646547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-01T05:19:21.647714Z digest=sha256:07dd1bb38ac7d405301afbe19d6a4b93bdf5ee725049911213f31bff91de71d8

Observation 2c37936e-c6c7-41b2-bcd1-c810c43c906f · inbound

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing cites this paper.

DisciplineGen-1M: A Large-Scale Dataset for Multidisciplinary Visual Generation and Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-03T15:48:35.117418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-03T15:38:36.388570Z digest=sha256:6cac6c75d9f30ab0fc25f13497be155a03ac1b4165504946786d53ce45973613

Observation fc9d60c5-a1f8-4468-a9f1-a5ba17bf8b9e · inbound

Vision as Unified Multimodal Generation cites this paper.

Vision as Unified Multimodal Generation SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-08T02:04:26.364613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-07-08T01:54:30.649092Z digest=sha256:ed74f19b2987a70f5cc3ebcce81ebc287aef3a06172ebd817a46ab011a859a61

Observation 9cace125-3dd5-41a2-88e9-ce4217fe4fae · inbound

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions cites this paper.

SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T23:43:35.072852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T23:43:35.072852Z digest=sha256:3081acd1f3d8e1d12ca569523f3c56c5b80a9dd906c128754a2b0aafa87a8638

Observation 4579cabf-9355-4505-80c7-f018bb3855eb · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T19:41:47.438584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:41:47.438584Z digest=sha256:e5ff7d13a00cab3bb13cee51b3492f85698f8010ac1eb2c505ecb4a717c05f75

Observation 7cb9eaa3-f7b0-4224-8e16-8cf6f3a294b7 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T04:17:41.402174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:17:41.402174Z digest=sha256:1f14bb28545ac101604b0257edf07f1ada84c61cac5d356612aa58b32c00effc

Observation 37a337e9-960c-47a3-aee4-a1388418953c · inbound

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs cites this paper.

STBridge: Shared-Target Alignment for Bridging Understanding and Generation in UMMs SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T18:56:57.619206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:56:57.619206Z digest=sha256:39031f7c6cbafbe09b23d269db5f767a56b89b348f9e16855a4474ff30c484ce

Observation 8c2401b0-516e-4693-a4ea-365964f7b622 · inbound

Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence cites this paper.

Unifying Generative Recall and Multi-Objective Ranking in a Single Decoder-Only Sequence SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T14:57:39.507916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:57:39.507916Z digest=sha256:49c7bc77c41c6ff5559b47297dd8d6529b4b09849b8e0ec077e702e5f3e33f34