Pith. sign in

Paper Citation Record · LEDGER

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models

As of 18 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2608.11907.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11907 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:31:06.104344Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7d4f1618-9d1b-431b-bca6-96194f70d497 · outbound

This paper cites Qwen3-VL Technical Report.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Qwen3-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.776988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.776988Z digest=sha256:24ea15c2fcc30ce7670e3028022a92028c45048c73c8c8123b21052501b31d9d

Observation 6a99321c-763c-4594-88a5-7081434e7ef9 · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.783381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.783381Z digest=sha256:0889aea81c6f4f584c06fe1be9d16849dc0e0e6578bb944ed4737e013a5fc8e1

Observation bdf9e0d6-4bef-412e-bde4-f77002b4f995 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.788627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.788627Z digest=sha256:ce448c9f1e3ef7f9285294006a05bb0ec6af46e63a3296e70989411afc2d4edf

Observation e5f19177-31d6-494b-a825-acfdffd0961e · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.794515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.794515Z digest=sha256:3aa64ecdf1f3c7fb8b2d5ff3ef8bd1c991845fbfa218c5c88b2a1a6206aab901

Observation 56198b8e-454f-4ee8-b399-46cda0b25a81 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.800105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.800105Z digest=sha256:97bb655a398a1cb57bd814db1d51a72b68fa05a9168a04fe149f41949d95553b

Observation 9bc1c440-1b91-4518-87fa-bc4479b3823c · outbound

This paper cites Agent AI: Surveying the Horizons of Multimodal Interaction.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Agent AI: Surveying the Horizons of Multimodal Interaction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.813474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.813474Z digest=sha256:3517e43e300bab194abad27bfe95f940445d70e353e1c795d50c12724fab1fc1

Observation 8fa3456b-b15c-4ffd-9df3-4b6d417627f5 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.819502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.819502Z digest=sha256:4462a3178cf4bad73a09c9c76851522591e677237485aed70cb07aaf1e4b8123

Observation c65aa755-dbdf-4fda-90ee-9b668d409b78 · outbound

This paper cites X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models X-Omni: Reinforcement Learning Makes Discrete Autoregressive Image Generative Models Great Again

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.824802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.824802Z digest=sha256:cc6a2f11a5899b35451d7aeaaba0b1db43c53b75ad06dafe9c3e1308955a210d

Observation 1bf49c15-62a3-4baa-b2f3-488e857f043e · outbound

This paper cites GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.830130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.830130Z digest=sha256:e17606793e8b11334af143c9e205cb4a48a1aa5031e9299051d8372ba71f7494

Observation 3d8a8f52-4260-4a14-b2bd-71a040360acd · outbound

This paper cites URL https://developers.googleblog.com/en/introducing-gemini-2-5- flash-image/.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models URL https://developers.googleblog.com/en/introducing-gemini-2-5- flash-image/

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:31:07.040660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:31:05.835396Z digest=sha256:87f761a9f140dc6962b451c4991a31b88268198148abafafd3be140414b0f4f8

Observation e1bbd716-2791-4ee4-8d65-46d2e49c0d52 · outbound

This paper cites CLIPScore: A Reference-free Evaluation Metric for Image Captioning.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models CLIPScore: A Reference-free Evaluation Metric for Image Captioning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.840495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.840495Z digest=sha256:aff729d584bfee173a89c3bfc926f1812c8ccf85917be8baaed1257d39adbcf0

Observation aa1f129a-d250-4cf5-b4ff-212837f7a3ff · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:31:07.024918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:31:05.845748Z digest=sha256:998fb988d69d77f29c832196bf462b1a82d4556455a8f5a304b72f52367c213d

Observation a7a6110f-2ffa-4444-a220-ef3b5b24f72a · outbound

This paper cites Gir-bench: Versatile benchmark for generating images with reasoning, 2025.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Gir-bench: Versatile benchmark for generating images with reasoning, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.851465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.851465Z digest=sha256:95e869c67f3f9fc9705fac556686ded3c349d7b0b9e76f3837b66dc1ae863cb5

Observation 976ebd1e-caf5-418f-a91b-14d336489526 · outbound

This paper cites Rover: Benchmarking reciprocal cross-modal reasoning for omnimodal generation, 2025.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Rover: Benchmarking reciprocal cross-modal reasoning for omnimodal generation, 2025

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.856548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.856548Z digest=sha256:07a5c6f260ba53328ff9bfd46ddc8b8b2417c6e16d4fc89615e853998e9aeb8f

Observation 4df0d804-feeb-49d2-a2c7-f9a0ceb95cbc · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.861660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.861660Z digest=sha256:46395a909d41d4f934ee7fc7f1bc188569e8af74303febe78340cf9648012ad1

Observation 1ff6ea94-dd9c-4d59-8c7b-8cbbbc56db89 · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Evaluating text-to-visual generation with image-to-text generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.866824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.866824Z digest=sha256:e453a4e6fc6ddfc98c21560aa33e4882e69533b6224a3036c5903d83d1603ba3

Observation d452d281-5483-4ef7-855a-0ac937c50a26 · outbound

This paper cites Flow Matching for Generative Modeling.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Flow Matching for Generative Modeling

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.871631Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.871631Z digest=sha256:751894d938b9211d17a24241d3ffc6186b88beaf91f09027499707048414b575

Observation c5ef1789-950d-4754-854c-b9eee19c0e37 · outbound

This paper cites Umnibench: Unified understand and generation model oriented omni-dimensional benchmark, 2025.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Umnibench: Unified understand and generation model oriented omni-dimensional benchmark, 2025

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-16T00:31:06.648930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:31:05.876038Z digest=sha256:a157c5f90238f93116f353f408653f2c07b1f1eba5b9ec132bf65aab374b8c7c

Observation f6b3b2ce-d380-4f45-b3e8-bc6a4c2f0844 · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models MMBench: Is Your Multi-modal Model an All-around Player?

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.880391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.880391Z digest=sha256:a22be212fe196967d32fd445f85839dea94f8ee8a54164a0c58d9d97d9cbab75

Observation 634f40eb-f42d-4f37-bd6f-3a400efd2a4f · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.886228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.886228Z digest=sha256:ae8368464dda050851f8baa4de6e1fa59763d3a32e0105ceced09aa76765d9dd

Observation 7d5da735-0e0b-4ca7-bb5d-94baec652197 · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Ocr-vqa: Visual question answering by reading text in images

Reference 21

Resolution
malformed identifier
no resolver link, observed 2026-08-16T00:31:05.891035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.891035Z digest=sha256:2e73dc84fb407d5f9c3847d5c5081d67645175c3b06805a794ff4811911e95d7

Observation 084b5dbb-a1bd-4634-b39e-14a46e56b3bc · outbound

This paper cites URL https://openai.com/index/image-generation-api/.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models URL https://openai.com/index/image-generation-api/

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:31:06.998583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:31:05.896089Z digest=sha256:99e1849ee6207c802f80808dc8f9a3e0011d5c3f4d493c38aebfd958a4b433fe

Observation 1ba0ab26-e5b7-4bfe-9877-9026269ec873 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Transfer between Modalities with MetaQueries

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.901086Z digest=sha256:1b7dfd6495e795a5ecf3d1c58667efd84445f7c74cdff1113e3af77159a14536

Observation 906ef5af-7494-42bd-b6a5-b738d3c492c4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.906338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.906338Z digest=sha256:7f526033df2ed17407bdf865f3fa4ba9fa93031d19ba40143bea2be6cf9d43df

Observation 3f9d7d50-5a78-435e-817c-0e09c284d389 · outbound

This paper cites Ovis-U1 Technical Report.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Ovis-U1 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.911595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.911595Z digest=sha256:633fe20f03b03447c74370b7861e60e9a6569501b6ae371ae088a81cd87b3629

Observation 05ab8237-b409-4307-aa26-3385b75ff41c · outbound

This paper cites From Understanding the World to Intervening in It: A Unified Multi-Scale Framework for Embodied Cognition.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models From Understanding the World to Intervening in It: A Unified Multi-Scale Framework for Embodied Cognition

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-16T00:31:06.407547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T00:31:05.916966Z digest=sha256:bfab688b10610c07916662dfcd03c763f5a56636a61e5b534afb4b0db9e5d4fc

Observation 6c900f6d-fd35-453a-8ed1-7a9d7f5d804e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Emu3: Next-Token Prediction is All You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.922909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.922909Z digest=sha256:5b6ef0bbb15a65193b822f928e63c216ce44243ff390667b6e652716973cef84

Observation 5460cad6-8b40-4e8a-a939-666cc619efa7 · outbound

This paper cites Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model, 2025.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Skywork unipic 2.0: Building kontext model with online rl for unified multimodal model, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.928340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.928340Z digest=sha256:c1bdacdd5bf5a3b93330372c0d99436dc594bd8c7ecfe1543da69e574ad0d7bf

Observation ce9a6e48-ea7b-492e-a043-a90dc897c898 · outbound

This paper cites Ggbench: A geometric generative reasoning benchmark for unified multimodal models, 2026.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Ggbench: A geometric generative reasoning benchmark for unified multimodal models, 2026

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.933256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.933256Z digest=sha256:1dd3938edec31209f2c3169d3c8093c36498584f9642fd4e72fdfd61879d28f0

Observation 859de7c2-1f64-40e3-811f-3ed850de4bd5 · outbound

This paper cites Qwen-Image Technical Report.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Qwen-Image Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.938087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.938087Z digest=sha256:1593ee36811ead41f22104e83cddbad810d4ebddcdfa81248e78b9ff43804296

Observation eb60004a-8d15-412e-a459-756ed89fada1 · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.943233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.943233Z digest=sha256:c9e58bae9582ef3354a8b90ec38208a02c48942b2476f45c116016bc25d2548f

Observation 9d58370c-8cad-437f-8316-962b99c4d3d8 · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Reconstruction Alignment Improves Unified Multimodal Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.948366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.948366Z digest=sha256:d411d08a91664ac12c012d8ae5a8a259ec9c25f8ad787971c5ad842efd973f04

Observation 949780f9-a252-4abc-9c0e-c6b7d2c12cfc · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Show-o2: Improved Native Unified Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.953814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.953814Z digest=sha256:33c8d39eb246210086c013312ec4b53eeb16f498b2de8381d6604c2b088fa2ee

Observation 44ea9f7d-e4d0-437a-b651-fdad456f97ae · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:05.959651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:05.959651Z digest=sha256:fbdf9d70022f18ad81d0d51505fcaadfde0f0257b315c0b2852454ee41112e8e

Observation 63581e84-093f-4b2d-87a3-a431e1addce8 · outbound

This paper cites Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Mmmu-pro: A more robust multi-discipline multimodal understanding benchmark

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:06.098205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:06.098205Z digest=sha256:5b6af91ea4588f001e024fd39dbbcc333c73cb280cf306a3b325747d635de4a3

Observation a6b62395-d16f-4434-8746-471fdc995a10 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Do You See What You Draw? A Semantic Closed-Loop Framework for Holistic Evaluation of Unified Multimodal Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:31:06.104344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:31:06.104344Z digest=sha256:67ede1d0bb7f2e4bff28b14abbf949ccee95896ed3f4201432b974cfc03f052b

Pith citing papers

No inbound Pith citation observations are available.