Pith. sign in

Paper Citation Record · LEDGER

Semantic Generative Tuning for Unified Multimodal Models

As of 21 August 2026, this Paper Citation Record lists 87 of 87 outbound references and 0 inbound Pith citation observations for arXiv:2605.18714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.18714 v2

Coverage vector

measured 87 of 87 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:31:10.578558Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

87 of 87 outbound references displayed

  • verified exact11
  • verified fuzzy37
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch35

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5069ae84-35d2-4dff-81c7-641275fd760d · outbound

This paper cites Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture.

Semantic Generative Tuning for Unified Multimodal Models Self-Supervised Learning from Images with a Joint-Embedding Predictive Architecture

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.289874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:074000112fb91f8d963ecb2343b5030b7049f1f2e300a7c9f56e20c2c390a997

Observation 91aa59ca-9b41-4098-a5ba-59ba06ebf9e5 · outbound

This paper cites Qwen2.5-VL Technical Report.

Semantic Generative Tuning for Unified Multimodal Models Qwen2.5-VL Technical Report

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.363633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:cc4892f78d54fb92959f8419f88ff3c2a62dd20c1e7301f448f91e4b1c587cf2

Observation 50058b55-5a95-47c3-a9e7-805a4b9b1cb3 · outbound

This paper cites an unresolved cited work.

Semantic Generative Tuning for Unified Multimodal Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-07-08T04:14:30.263200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c807e8f0e806a897afb7957e6d5f64d781bc50d9fd717c5cb630500564030c5b

Observation 9bd1b411-d7b2-4a7b-8cf0-70f8a83cff34 · outbound

This paper cites an unresolved cited work.

Semantic Generative Tuning for Unified Multimodal Models Unresolved cited work

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.294989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c21687a88b6ad7bf0ca42ee974fbd9c7a75102e860cbc4b6a94abd8e46cc9410

Observation 6c54261b-985b-4f7d-b622-02a9b1fcc97f · outbound

This paper cites an unresolved cited work.

Semantic Generative Tuning for Unified Multimodal Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-07-08T04:14:30.258471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:2840e770d788c09a5514fd0e51190c9bb3b1241294128b0be1b75ab688558d3b

Observation 4eef1f10-0a3b-45b9-b84c-133230ebea62 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Semantic Generative Tuning for Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.297856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:317d7a48d1376a19b92da4dfd96a69e1e36e209aa68862df9e06eedba03117e9

Observation 1ef9940f-ad48-4702-9956-cea6d8752e70 · outbound

This paper cites Deconstructing Denoising Diffusion Models for Self-Supervised Learning.

Semantic Generative Tuning for Unified Multimodal Models Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.285093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ec6b3f37775ce7774cca29279ca419e47d7eda1e5526a0f5c0c083f15ec4da8e

Observation 9fb063e7-fa70-4bb8-9149-d8de6808c76c · outbound

This paper cites NeurIPS36, 49250–49267 (2023) 2.

Semantic Generative Tuning for Unified Multimodal Models NeurIPS36, 49250–49267 (2023) 2

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.273323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:fb087a6e5ac0849c37db31d02e1610df8237b90c8282e7e9484391aab528e7fc

Observation 00debac7-6a8f-423c-89fe-07c3077de651 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Semantic Generative Tuning for Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.287454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:4cf4d3571572f7675349ce6cb375688b936ebab205617794dd0bac290c35490d

Observation 50c4e7d4-b29d-4958-ae7a-df2105fd6590 · outbound

This paper cites In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2.

Semantic Generative Tuning for Unified Multimodal Models In: ICLR (2024),https://openreview.net/forum? id=y01KGvd9Bw2

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.263481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5145ee1d070a650eae6760916b8618991fe98672e014ad140d54f556f7f44615

Observation 93a8e812-31d8-42ea-a74e-ea7618f38ad6 · outbound

This paper cites arXiv preprint arXiv:2511.23386 (2025) 4, 7, 9.

Semantic Generative Tuning for Unified Multimodal Models arXiv preprint arXiv:2511.23386 (2025) 4, 7, 9

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.361216Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:db25b329dd22847889a1720233966cced7dbced9ed2daeaac040c5b9472aaa2f

Observation fbcbc4b1-fa1a-46cf-b12b-53630a89187c · outbound

This paper cites In: ACMMM.

Semantic Generative Tuning for Unified Multimodal Models In: ACMMM

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.264936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:24cbdc52140a90f1925cf3d775a3427960f7859459acc1d7b815c2873433e9b0

Observation 1fe06d48-a184-4ccb-a1be-4b42f43d4af2 · outbound

This paper cites In: ICML (2024) 1, 4.

Semantic Generative Tuning for Unified Multimodal Models In: ICML (2024) 1, 4

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.275464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ca8e97a403887551282ac9ac941b4e31ccb4aacf84aa483eebfc8248d5e702f6

Observation aabe6c3f-09a7-4701-9aeb-44f0e2593404 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Semantic Generative Tuning for Unified Multimodal Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.356370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f6987d064736538dabd0af3339c121dfa5747a72232d70a947369b8889c63d5b

Observation c5943552-c23f-4eb4-971b-40db9737fc5e · outbound

This paper cites In: ECCV.

Semantic Generative Tuning for Unified Multimodal Models In: ECCV

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.265114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:745f68723ee2a0b98608999b62e9f6df42ddcefb091349422ccd6e14f8008317

Observation 41f4ffe9-b6b8-4ff8-819f-eb9b5000c5c0 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Semantic Generative Tuning for Unified Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.371150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c47c67ec6ef695577a55eada76bc8b33bacefc4a24b6436bb3c9f89a9da7b80e

Observation 6be2cc4f-d226-4388-bd58-8cfb49e136e9 · outbound

This paper cites Diffusion Models and Representation Learning: A Survey.

Semantic Generative Tuning for Unified Multimodal Models Diffusion Models and Representation Learning: A Survey

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.368536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5bafa577aa676ec06a8fdad1d2c66d8b6025e6b43f87ce2fa11b980f5f5bba00

Observation 0c937d14-e652-4430-8229-1401fa97dec9 · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

Semantic Generative Tuning for Unified Multimodal Models SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.298289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:1f0c0c29b605643240b27f6836d589ad46b4e84d00e05aac829d56659cdc7e88

Observation 2f9793e8-a9b8-4c72-bee4-7e3e634250c5 · outbound

This paper cites NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14.

Semantic Generative Tuning for Unified Multimodal Models NeurIPS36, 52132–52152 (2023) 3, 6, 9, 14

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.302618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:a75f3933026e306bad76993eee647501d25bcc0c4ddad24a16447600ddfcc709

Observation e4e1299a-a22d-4efe-b8d6-df67e330df29 · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.250657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:3464ddc4bfe6b653b3ebe8db21261edec830ec987e885fd8a6544999e47d2898

Observation 2655f109-2f6b-46f6-8851-49aede8689ba · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.252655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:805f16e3aafd56ae76001a20670df2e8883a6a3f47c5813aa1e4864d8271d123

Observation 0b1ee9d5-d44f-4a7d-ae15-25675dfab6fc · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.266566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f1da1a4c8c8706021cc043fef6e9913f153ad65de0d99fef8c1e694ddcc3ae4a

Observation fd3e675b-41bc-41f8-ae7e-891f95ba89cd · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.261565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:3408b7688ea1c7d21affa439c9fa0360688e09151b46072ca7cd0e0cfdd9782d

Observation 5654b223-4bc1-4bd7-8e81-9d7232d8ee6b · outbound

This paper cites arXiv:2402.03161 (2024) 2.

Semantic Generative Tuning for Unified Multimodal Models arXiv:2402.03161 (2024) 2

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.300676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c2fce09354b8748dbf921c07b5d5e6dbd9524206329bf73b100ea9a36cd7538f

Observation 11739eda-40de-472a-8834-49515769321f · outbound

This paper cites Unified language-vision pretraining in llm with dynamic discrete visual tokenization.

Semantic Generative Tuning for Unified Multimodal Models Unified language-vision pretraining in llm with dynamic discrete visual tokenization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.303308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:d20b2f718b7462caf5182a2f419b29603aaa0bee6084677a34540a52b7b75747

Observation 8c6ca92d-2d4c-4982-8055-a94ba4c5d3c2 · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.279771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:37217cf0ab6df1dccfe10dd0dda593cb4475ced6464fc3d09cf71ed1556f353f

Observation 3d1056ca-2311-4c08-8f49-2f37ead29be4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Semantic Generative Tuning for Unified Multimodal Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 27

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.292872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:9b8e100bde04733eb8e3ee4c3e44432b5b605970880499ab7e191f11f0d36bec

Observation 919769fd-70e1-43cf-b41b-101e1a032f0c · outbound

This paper cites In: EMNLP.

Semantic Generative Tuning for Unified Multimodal Models In: EMNLP

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.300286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:100014a3c3c79551ae687f5ef24c16b121650c43264990ce305d5bec061c2c8a

Observation c5a50c4c-7201-4940-9b01-6ab492e60428 · outbound

This paper cites Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback.

Semantic Generative Tuning for Unified Multimodal Models Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.303188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ca07d491f4c55c92647c12c93962491ff9c9d51db583013966264f95d29a31c6

Observation 9b00b86f-3823-48db-91b6-57e305284d78 · outbound

This paper cites arXiv:2512.19680 (2025) 5.

Semantic Generative Tuning for Unified Multimodal Models arXiv:2512.19680 (2025) 5

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.314100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:1a035a92099cfd2655315172d0bf0b44017e5198a9d6b758c110fb11a701594d

Observation 2839ba80-d757-4bf9-bea9-83025e038f43 · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.349052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ecb23f111ec79f8b4ca44c54407490f330202e3ea4e166a725cc50013cab6a71

Observation eca460a0-1964-4681-968a-b355346cdfb5 · outbound

This paper cites Transactions of the Association for Computational Linguistics (2023) 6, 9, 17.

Semantic Generative Tuning for Unified Multimodal Models Transactions of the Association for Computational Linguistics (2023) 6, 9, 17

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.284407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:10497f98ec2a086efff56813404c03a19c9a0c90122f365678d5dd5c62ce81d9

Observation 8b85ce76-4af7-42eb-8858-080038a20a40 · outbound

This paper cites In: NeurIPS (2023) 1.

Semantic Generative Tuning for Unified Multimodal Models In: NeurIPS (2023) 1

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.281513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:990b115280309b1fc1ca4189add4700d966a65f2a2e740bfad2d5733000ca228

Observation cfd9a9a7-8c76-495d-8023-99e220f2c01a · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Semantic Generative Tuning for Unified Multimodal Models Step1X-Edit: A Practical Framework for General Image Editing

Reference 34

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.279776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c16a2fb31561d131de7423fa76a36958d4739885bf883e3511e79766ac110d54

Observation 91889e4e-ecbe-4b5e-9df0-b70ee26b972a · outbound

This paper cites an unresolved cited work.

Semantic Generative Tuning for Unified Multimodal Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-07-08T04:14:30.271914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:580ce939c8f8aae34caccd66ebc4e3e7b94b9c23a0f455500bc98fbee0528d9c

Observation d771c448-949f-449d-bbb7-be3e0adf706f · outbound

This paper cites Science China Information Sciences67(12), 220102 (2024) 6, 17.

Semantic Generative Tuning for Unified Multimodal Models Science China Information Sciences67(12), 220102 (2024) 6, 17

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.270090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:bb96614c0b407d1f1773a629d42904404240a009e97584594b610db6104333c9

Observation b5181804-620d-4b7b-a402-dd4c8fec86f9 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

Semantic Generative Tuning for Unified Multimodal Models Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.290588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:0d370dfa2be4cee0a9c417f31c65938b7b0858d857ffa56d66b4eb6388525755

Observation ba93729d-fc8a-434e-bd9a-28425fa19229 · outbound

This paper cites In: ICLR (2024) 6, 9, 17.

Semantic Generative Tuning for Unified Multimodal Models In: ICLR (2024) 6, 9, 17

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.279430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:654b2f65d13feafe11a04f863d68b25655568b72b38d1338222383e25d24ed17

Observation a17107f4-4c18-4c00-b75d-ae39e883c220 · outbound

This paper cites In: NeurIPS (2022) 6, 17.

Semantic Generative Tuning for Unified Multimodal Models In: NeurIPS (2022) 6, 17

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.277392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:40284de9d245b8944b0f03c2f3893e4dfe660ceebf15636fa9b931940d6258e9

Observation 1545c2a8-dd4f-4398-8250-d4b01290cf60 · outbound

This paper cites DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception.

Semantic Generative Tuning for Unified Multimodal Models DEEM: Diffusion Models Serve as the Eyes of Large Language Models for Image Perception

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.271961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:73989181a411cabe8125fcf4a4f20089657ce8a18526d8e7047eb4af1b9da6d5

Observation ef58fd2b-aa4f-4fb8-9fdb-218fdb60d15e · outbound

This paper cites Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a.

Semantic Generative Tuning for Unified Multimodal Models Unitok: A unified tokenizer for visual generation and understanding.arXiv preprint arXiv:2502.20321, 2025a

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.274664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:82292b71323a7f17945dfee1e6dd2f7f09905093930285a735feb06743bba5bd

Observation 2811b71f-04cf-4dce-903a-486beb3ba0cb · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.239596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:8e5c318ea308e6c96685039fe31e7420e9918513dfa1a3a18d4b11a3a9f8a61c

Observation 42b56341-4574-4b54-a5ee-59f32cf6cfd0 · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.268126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:6e47770a5a694e5448ac848a2b010077b77d960edc73eae6229cdbcbe8a33094

Observation 89db8fce-8641-448d-b60c-d4b3cde6d150 · outbound

This paper cites In: WACV.

Semantic Generative Tuning for Unified Multimodal Models In: WACV

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.232102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:84cb8810b72ff05a0652233de2347ffe85903c349cecc3917088c4c8e505309b

Observation 18951bd3-bbf1-4574-bac6-89c2e71c7feb · outbound

This paper cites WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation.

Semantic Generative Tuning for Unified Multimodal Models WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.266110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:b4835fdf86acda9e68efa51e1377f57e161fff0404a24574f7209b8ba45392a0

Observation 30c633d1-3a2e-45fe-b414-7ebcad8ff710 · outbound

This paper cites Transfer between Modalities with MetaQueries.

Semantic Generative Tuning for Unified Multimodal Models Transfer between Modalities with MetaQueries

Reference 47

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.277120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ec9ac5217b55df76b9032c77b46861fff8a42ef3b9e0ed9f656e486c39793059

Observation 5daf2b32-b558-496c-8b53-bd81d9bdc28e · outbound

This paper cites In: ECCV.

Semantic Generative Tuning for Unified Multimodal Models In: ECCV

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.256144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c181661f85d7a1a79c3c628bf5c00160f43d3b6a56fa4080fe75a9a378dd132f

Observation 3dcad568-6cbc-45ae-ba83-ae83dd0f460c · outbound

This paper cites In: CVPR (2022) 14.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR (2022) 14

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.252462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:20f155750765ab16a97e22c22bf283df2d28bbba85e02ce1fb6baf6520c6190a

Observation a64ca134-4288-40a2-9b15-810d901ff791 · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.303767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:32f85fb41baf9adf652f96028bc295e60519452fcace9a6ac77996ea6f2f0d68

Observation e554ff05-2944-404f-b193-242c8dc4740e · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Semantic Generative Tuning for Unified Multimodal Models LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.282413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:2c8f2df8f1b4d4f983bdb8985e32672d83a2f5125a6c2b4ca9d9d0c3f5c686e5

Observation b5df93c4-7775-4fc3-96df-d1766ddd1863 · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.299134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:7f616733e40b542c33ad37e0d31bd37064a64be85c3c4248f533e764af043120

Observation c307a630-d72c-4e80-84f3-a5d883a7147a · outbound

This paper cites Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation.

Semantic Generative Tuning for Unified Multimodal Models Generation Enhances Understanding in Unified Multimodal Models via Multi-Representation Generation

Reference 53

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.355652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:612fa42f9b4ee5a79a7b5de638ffab2c8b2d3a500a5f555a514f9aa6bde3191a

Observation 0261830c-42d7-494b-90e0-5a0aeca40ffb · outbound

This paper cites Unilip: Adapting clip for unified multimodal understanding, generation and editing.

Semantic Generative Tuning for Unified Multimodal Models Unilip: Adapting clip for unified multimodal understanding, generation and editing

Reference 54

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.360536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:11486977f8288e9031b5ddbb11a7da87984abf8ea27b22fb0a3864d9e76a43bc

Observation 9ad11093-71ac-4cc7-9e69-9ec7a9d0c31d · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Semantic Generative Tuning for Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.366165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:41e13d6b1d748ce46fba728e058444b38aaa4c4105216d5ae95e56f82172f3c2

Observation 82e477aa-c3f3-4ae1-8705-567a68ddeaa0 · outbound

This paper cites NeurIPS37, 84839–84865 (2024) 1.

Semantic Generative Tuning for Unified Multimodal Models NeurIPS37, 84839–84865 (2024) 1

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.285295Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5fe19d976e556c106907d7c3ec5c1dc24bb539b508425ae74e77074725ef07f7

Observation 5ef342fd-00f5-4d19-91bd-4ac4e6573c01 · outbound

This paper cites NeurIPS 36, 48382–48402 (2023) 4.

Semantic Generative Tuning for Unified Multimodal Models NeurIPS 36, 48382–48402 (2023) 4

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.296756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5cb154c0375009818ed0b1ea9e8285539055a3545ede6d133fd78d20bdf1f182

Observation cbfe9079-fe0d-4752-83ae-72b4278ae101 · outbound

This paper cites NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17.

Semantic Generative Tuning for Unified Multimodal Models NeurIPs37, 87310–87356 (2024) 3, 6, 9, 11, 17

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.292504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f91b43c09737e974d1d5b574e337f37eafb7b5abf8fca3ec6637d53c58f6f728

Observation 332531c5-633b-4628-936e-42514fec5c9d · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

Semantic Generative Tuning for Unified Multimodal Models MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 59

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.351343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:f1082d3b704de2a4dc6b96075e1caff756728b6f0f343d3b05dfdece539be56d

Observation 4c3d41c8-b3ec-4d8c-b412-697f22fc8266 · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.290669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:3f638d0c24a129257fc526784b5ed2693d7821f9e109009fb72e38d73519f5a4

Observation fb0388fb-fd3c-4570-b15b-345d8c0f86fc · outbound

This paper cites Reconstructive Visual Instruction Tuning.

Semantic Generative Tuning for Unified Multimodal Models Reconstructive Visual Instruction Tuning

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.342775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c16c841a0032256574ba07d88da1f4bd487c8889cf70c10f00011ca68287bfcc

Observation 76a3f3d9-c181-4765-9b82-cd61480e4c65 · outbound

This paper cites Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models Skywork UniPic: Unified Autoregressive Modeling for Visual Understanding and Generation

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.339365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:7a5aa7dd497800ac538e34308045524b3d3d4b50f4e43edfec5aafb2c5a1aba0

Observation 7a27550d-7865-4f7b-ba42-294aebb97ed5 · outbound

This paper cites Diffusion Feedback Helps CLIP See Better.

Semantic Generative Tuning for Unified Multimodal Models Diffusion Feedback Helps CLIP See Better

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.334765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:8460236192406f2747031208b5a61c33de33472d85b689c0588b549ce0eb7c67

Observation 6db027a5-7b2d-4d44-8a1d-acdb3f2d4181 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Semantic Generative Tuning for Unified Multimodal Models Emu3: Next-Token Prediction is All You Need

Reference 64

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.341727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:cc33b9a64038c05bbba6ad68c7eae616db44ca4dcced887de3bb8db959752df8

Observation 6e16c4fa-fdc6-40ab-9284-7fb115693fed · outbound

This paper cites In: ICML.

Semantic Generative Tuning for Unified Multimodal Models In: ICML

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.295346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:a103a3cf8b063175620f50dc75f691e48ae534248f9a4f1e76faada63278b8c0

Observation d1b55d7b-7316-44e3-8f0f-ad044b2949b3 · outbound

This paper cites arXiv:2510.22946 (2025) 3.

Semantic Generative Tuning for Unified Multimodal Models arXiv:2510.22946 (2025) 3

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.344146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:9d1f8b40445b2664574cffc118641e5b368d6216e0a6604fd056c8f29a0fd3e6

Observation e38b4d08-ce74-4295-9a42-f782bd6c3086 · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.301526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:06b5fb1ea64d166799cd086a1c251cead39aefc8c9a9f52680f36eeb9c6fbfa7

Observation 6996a4a6-69bc-43af-8a25-558fcbb8b5c6 · outbound

This paper cites In: ECCV.

Semantic Generative Tuning for Unified Multimodal Models In: ECCV

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.289230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:fb9f9dad153aab851cb9deb1715fb10c24ae83d959b0f0323f79948a96e35d77

Observation b9f590d8-3066-4f3f-8eb1-2c565f935d8e · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 69

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:35:00.331203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:c1df48c0da271da26d631cdfc803e19e385aa26870fad83c4c4ca806ee85336e

Observation 1f5fa638-0c5d-423d-9908-11978ccb5b5f · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Semantic Generative Tuning for Unified Multimodal Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 70

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.328081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5378e962854ae350b72f28949031a90b8cfa40fd8a4702b39e9f26d7496255dc

Observation 0a6ca79c-15e4-496a-bd83-0a18458ad585 · outbound

This paper cites Representation entanglement for generation: Training diffusion transformers is much easier than you think.

Semantic Generative Tuning for Unified Multimodal Models Representation entanglement for generation: Training diffusion transformers is much easier than you think

Reference 71

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.322082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:a9efeccb77e64a25df270b94eaf35c4d3cd3199186f264aeb844fd760ae5d1b5

Observation 6f178c0a-3fa9-4ffb-b464-d98cfe592b52 · outbound

This paper cites IJCV (2025) 3.

Semantic Generative Tuning for Unified Multimodal Models IJCV (2025) 3

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.298028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:554084bb897039759cb39184f450cd9dcf8bc625e696a66e27fe98930c8a2b56

Observation a318c7ae-ef0e-480b-9093-0fd3fadac828 · outbound

This paper cites OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.345303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:8a48f310fcad1b82673d47d8854cd76295579fe59c0e724b0c1d91143d3695b2

Observation 905b2310-02e7-442c-a011-1f3afbc44737 · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.305889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:8d21a38460cc7d6eaf9798ff3381588fc556a75796a162b5009605fb4a2a0e81

Observation 63c989ec-88ac-4c96-a3bc-3557f29a6015 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.358106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:68bc4e263a29807b92d8ea6e7323f35a1e42d30506a11682e8aaa3a006cb4452

Observation 81569916-f0de-4b9e-8010-4a5dfc93a542 · outbound

This paper cites an unresolved cited work.

Semantic Generative Tuning for Unified Multimodal Models Unresolved cited work

Reference 76

Resolution
unresolved
raw_fallback, observed 2026-07-08T04:14:30.277719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:b3255e07da93ac8354b788e9f2dceac7f0b99a07490192b4480503edb8b59981

Observation 11442d7f-904d-45d4-b420-4be021c9bb37 · outbound

This paper cites Reconstruction Alignment Improves Unified Multimodal Models.

Semantic Generative Tuning for Unified Multimodal Models Reconstruction Alignment Improves Unified Multimodal Models

Reference 77

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.324906Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:1ba481431071048e73d24c0e2d4f9e946471418d855ae8e23377c549cc924131

Observation 3aba8f9c-2e70-4279-96b9-adf0d88927fe · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Semantic Generative Tuning for Unified Multimodal Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.334115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:6f385b696002d846ec111491777ac0948aa8ec199b88352e3319f8658ab6b89d

Observation eace8d0e-30cc-40a4-a234-345f7711d286 · outbound

This paper cites Show-o2: Improved Native Unified Multimodal Models.

Semantic Generative Tuning for Unified Multimodal Models Show-o2: Improved Native Unified Multimodal Models

Reference 79

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.346694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:4e9fa3a5185cfa5462c0418b7677b161ad873a10a81441a34eb7121292a830a7

Observation 8f267e0c-7643-432f-907c-d37217d57dac · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.283453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:ceaab3af778f358065051ac9ae047922fa2918d9a50f746004fbf8c061a17935

Observation 4db69def-9e93-4ef9-9ae6-279de0fcc139 · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.268417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:3b07e9fcb00063e490cb984f1b48abf9d515b03a250c308c3870501986895ae3

Observation 64c7a105-3c96-4fc3-87cd-7290e11ba0ff · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think.

Semantic Generative Tuning for Unified Multimodal Models Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.353982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:e9698df07ebd97c8e6e22218d4f6cb9c9cc6cc86381819a7d149c47b509ea21b

Observation bbe24a51-eb74-4d76-8708-9f024ca66c49 · outbound

This paper cites arXiv:2509.18905 (2025) 6, 9, 17.

Semantic Generative Tuning for Unified Multimodal Models arXiv:2509.18905 (2025) 6, 9, 17

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:35:00.358770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:cb412442de3508933de0e12c61bcd1c62f512e0e1aea6600fc843c026dd4f291

Observation f7b1ceee-7580-4591-8ed1-ee6e72c5580b · outbound

This paper cites In: CVPR.

Semantic Generative Tuning for Unified Multimodal Models In: CVPR

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.269781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:5b08bb7e80640ad9d821d655e7a883967af7c351f36e2c38109e052f74cd72eb

Observation 141c38d5-11d8-45e6-acf5-626c0776179a · outbound

This paper cites In: ICCV.

Semantic Generative Tuning for Unified Multimodal Models In: ICCV

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T04:14:30.307437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:65f47c36b2617c3dff0ddd758cdc9147f5792daa9db0aa509e7fc8a9cbceff5e

Observation 69715a54-940a-4a56-bd7d-4de3a12c796a · outbound

This paper cites MLLMs are Deeply Affected by Modality Bias.

Semantic Generative Tuning for Unified Multimodal Models MLLMs are Deeply Affected by Modality Bias

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.347720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:cf4e2e813e075bf3798f467812aa4df6c5b148fe5ba0798d6795abe01def2960

Observation 74b39d7d-3100-426d-b216-1338aeb053c0 · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Semantic Generative Tuning for Unified Multimodal Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.365861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:d16ff02d46bf18a70ce328c2344f3b09e1814657ea72b46fac6e027ad92fc10d

Observation 061a9d94-8747-45e3-a2a1-d89cc9b3bf21 · outbound

This paper cites VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model.

Semantic Generative Tuning for Unified Multimodal Models VARGPT: Unified Understanding and Generation in a Visual Autoregressive Multimodal Large Language Model

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:35:00.311433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:9988b2fa446b4c0b3167eb264249c99831db94399f21251c72d90ba2706b32cb

Pith citing papers

No inbound Pith citation observations are available.