Pith. sign in

Paper Citation Record · LEDGER

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

As of 12 August 2026, this Paper Citation Record lists 100 of 113 outbound references and 9 inbound Pith citation observations for arXiv:2412.09604.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.09604 v1

Coverage vector

measured 100 of 113 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T17:00:09.097663Z

measured 109 of 109 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:35:40.295495Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T05:50:24.280970Z

Reference resolution

100 of 113 outbound references displayed

  • verified exact1
  • verified fuzzy12
  • unresolved87
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 64717b3e-5766-41ef-bc33-24fb9ff5f58d · outbound

This paper cites Qwen Technical Report.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.566199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.566199Z digest=sha256:8a421b7393e70e7db6bf3f69474aefe1e3b60b00a26b2eca6c5424ac56df56d3

Observation 52df0d8e-5e4b-4c91-ad81-67c963194a47 · outbound

This paper cites Introducing our multimodal models, 2023.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Introducing our multimodal models, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.572203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.572203Z digest=sha256:cb52e4a9733e254e293bc05fe4c653da105c6b025bb986460da2e3955af804d3

Observation c06b6b5b-7756-4c95-8d47-81284370bb54 · outbound

This paper cites Improving image generation with better captions.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Improving image generation with better captions

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.577193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.577193Z digest=sha256:93b7088cdc4faaf9c4b07e8e8cb8927d83d1303bc615741bbcc1a6813a48d6fc

Observation 74690afe-ae2d-4c07-a240-ce62d8a052d7 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding PaliGemma: A versatile 3B VLM for transfer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.582788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.582788Z digest=sha256:538ad1ed444a6832cead6c2703bf2bdc288416f193183c95acf3cbecd986cb24

Observation 9b31bf2f-9ede-4edf-889e-4b218696c7d6 · outbound

This paper cites Scene text visual question answering.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Scene text visual question answering

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.587261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.587261Z digest=sha256:f39c44d72d332cf220503c4322a96b23565deb7d78abc19aee57a939754da814

Observation 149636dc-4047-4c18-baf0-d8e3ccd562e8 · outbound

This paper cites Coyo-700m: Image-text pair dataset.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Coyo-700m: Image-text pair dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.591759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.591759Z digest=sha256:bd2ec77a3c285a3d7086be5fdb3a41df1309e7f4dfa26f3227a886a6290df099

Observation 469a3509-912e-4cef-ad0f-e7775866a0cc · outbound

This paper cites InternLM2 Technical Report.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding InternLM2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.596644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.596644Z digest=sha256:8a64bffafa13852d35ed24b84147b58a7fa89c579feb620b4442fa1bb4fc2ada

Observation 2c8501eb-da92-42bd-9a93-275274bfe0d2 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.602001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.602001Z digest=sha256:37910b5b5ab37e9ff5931e3f4c6585256af83f7d845f4139cc3a3562fcbf26af

Observation 8fa3bc14-ee38-46f5-805c-451a5aa2d819 · outbound

This paper cites PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.607166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.607166Z digest=sha256:9b537ffd04692e203d54f469d363b27f89a7646087afe2f9224637f2bcaf03ce

Observation b0c3328d-b256-43ab-91ea-11c8c42329ab · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.613191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.613191Z digest=sha256:710f6dde37bd3e3eb719919a648c77ba9d2a731c8bb5b749ffa259d95e3844d1

Observation 6eaa2e8e-9a92-4ffc-a583-6cfb99c13621 · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.618440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.618440Z digest=sha256:79d44912061dfa6f642f7c98afd1ef5f41809b334dba349703ea83efaf567b20

Observation ff0bc6a6-0d81-449a-b057-7d09068f27b0 · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gonzalez, Ion Stoica, and Eric P

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.623487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.623487Z digest=sha256:b52441e4d38208803baa972f4518db93cbb33a4c5d579512580b71ad23398ed2

Observation 4ae26fac-b719-49c0-bf2e-634442dfe42f · outbound

This paper cites Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar2019 robust read- ing challenge on arbitrary-shaped text-rrc-art

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.628478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.628478Z digest=sha256:cac63a7ee9cb7561f988777d28578dcc2c5ab81c250d66833d02191e6c303a68

Observation dcd04dcb-f70d-4cd8-85c7-3ee7375efbf0 · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.634317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.634317Z digest=sha256:3c0f15c63a8492791621f355a00c82bf64956203a9b9ccd031ea69764c88374d

Observation 38b1d0b2-bff2-44c0-9818-d6b0adaaead3 · outbound

This paper cites Simple and effec- tive multi-paragraph reading comprehension.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Simple and effec- tive multi-paragraph reading comprehension

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.639363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.639363Z digest=sha256:e759fdaa0948c2324e93a9ea1f91ffc35a866a29ca336cda370e0e0f765ce79c

Observation 335d9490-860c-4909-9367-94d530bf6ff9 · outbound

This paper cites Opencompass: A universal evaluation plat- form for foundation models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Opencompass: A universal evaluation plat- form for foundation models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.644117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.644117Z digest=sha256:bb59575374007073ab0fd2374726fce6ff9be3b781d82ffb43e9d6ed86cc9913

Observation 5d232235-38d6-4cd3-8067-9e0d3da90210 · outbound

This paper cites Funnel-transformer: Filtering out sequential redundancy for efficient language processing.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Funnel-transformer: Filtering out sequential redundancy for efficient language processing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.651433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.651433Z digest=sha256:412b883f1509ac3640d35f37c8461e165be3b9d88f0e1841742a88f26e685727

Observation c2dc95c6-0a66-4cec-90c5-5d0ba44fe7e3 · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.656490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.656490Z digest=sha256:c149e306b575f51cd562fdabb481823e1872b67d3284e33c1c979ee442a0500a

Observation 16ee3840-c0b6-4cfc-b6cd-a81b11080e4d · outbound

This paper cites Unveiling Encoder-Free Vision-Language Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unveiling Encoder-Free Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.661560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.661560Z digest=sha256:362930452f7bdf0babcbd48548d3134c0bd0187d0ca0b37edd964aa6c9a87724

Observation 75c0af2d-1227-46a2-ac31-40fded98835c · outbound

This paper cites Dreamllm: Synergistic multimodal com- prehension and creation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Dreamllm: Synergistic multimodal com- prehension and creation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.667730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.667730Z digest=sha256:0252df0bd7c84b43fbcbc5646e59148cbe9706b96f12cf53bb6ef9269801fae0

Observation 911769e8-a758-46ad-a81a-229f218f1da3 · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2024.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Vlmevalkit: An open-source toolkit for evaluating large multi-modality models, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.673408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.673408Z digest=sha256:0c18f48499552e634fedcfb5257e97bf0f1ec51a1836b20793b99d51f8815937

Observation f1a80be6-4fd1-4f70-b0ce-c21512933a8e · outbound

This paper cites Taming transformers for high-resolution image synthesis.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Taming transformers for high-resolution image synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.680475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.680475Z digest=sha256:2d81432787232a64b9aba5efd09263e22e7dbdf392dacf5fc3de44cc0a02d40c

Observation 013fd86f-f901-4e86-9735-ee1bfb0fec5d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.687160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.687160Z digest=sha256:18d7bd8f9dd7310ba43b3279481fd2e5e4eb74c26beb6c3f7d98118e5fe671cc

Observation 68bbb49b-9e7c-446d-8a3d-0c96b5ba8c6d · outbound

This paper cites SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SEED-X: Multimodal Models with Unified Multi-granularity Comprehension and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.692902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.692902Z digest=sha256:d092b849933d4e4d0868608e7b5788966825d7b5790c1d4582321838d3a26fd0

Observation 04779d43-a551-4012-8363-8c79d7b22d1b · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Geneval: An object-focused framework for evaluating text- to-image alignment

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.698351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.698351Z digest=sha256:96c81945cb36f62291c216abe6e7782db8bc0b8c2b4b0f9b3f89be6e3588a374

Observation e38d3f0f-f276-4cc4-be6f-7748058118c1 · outbound

This paper cites Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Wukong: A 100 million large-scale chinese cross-modal pre-training benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.705267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.705267Z digest=sha256:06caef3227ee0f7dd171a9f0dd5e678f33ba4e31f6a82456f00afe5dbc6b23fa

Observation 9be28103-c8dc-44c3-b7b3-8d19d339168e · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.710316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.710316Z digest=sha256:f9f3f0c556697490533b7097e07425bd7909db263a25e4e7bf24c4bf1064e274

Observation 106f2a8d-9f05-48fc-8eb2-2614b2393266 · outbound

This paper cites Block Transformer: Global-to-Local Language Modeling for Fast Inference.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Block Transformer: Global-to-Local Language Modeling for Fast Inference

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.715269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.715269Z digest=sha256:c39fe03d57fb50b742e72815a6d51b7951dfc7218a2686316242956d04d8d6e7

Observation 9dd45ea8-1df6-4c6e-bd3b-0d8003951a46 · outbound

This paper cites Hudson and Christopher D.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hudson and Christopher D

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.721507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.721507Z digest=sha256:e1b2305f3e9e50ed9bb138c0c1a6d17b366819ce4367c61f67d5d07d1697b4cd

Observation b620bf40-8876-43f4-9941-598a7bd4da2d · outbound

This paper cites Unified language-vision pre- training in llm with dynamic discrete visual tokenization.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unified language-vision pre- training in llm with dynamic discrete visual tokenization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.726573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.726573Z digest=sha256:fe1e6fcb681eeebc45aaac4d46fa6c56c77ccc0830cb86563035ba5fa7cb4cc8

Observation 272e9666-5abd-41ee-ad29-07d525c51c4a · outbound

This paper cites A diagram is worth a dozen images.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding A diagram is worth a dozen images

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.734219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.734219Z digest=sha256:fc149c210ee8cb0e26d11efbc3ec1192d8043ee748c953df755439e753e49df6

Observation 498a6a59-1ea5-4cb2-aea0-a32e5f9a3bfd · outbound

This paper cites Ocr- free document understanding transformer.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Ocr- free document understanding transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.740322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.740322Z digest=sha256:87e6790b20a436fa3d91b29969ec457ff9157d92a039a1bd408c6b487ba00f58

Observation 2b3bda50-48e1-46a2-83de-6c43b090a86a · outbound

This paper cites Segment Anything.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Segment Anything

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.745540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.745540Z digest=sha256:0351fbeefedd81d71dbbdd24bd453833804127d9b61729d23e54e6dc3dcd728b

Observation 9c8051d8-123f-4bc3-8092-360ddb98fc38 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.752892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.752892Z digest=sha256:595411dec82b99eb647f70a73e724691d8a9e757b3395ac8e80bb7084a0e3876

Observation 24f69ec6-b6c5-49d2-ae41-769e4211d4db · outbound

This paper cites Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Playground v2.5: Three Insights towards Enhancing Aesthetic Quality in Text-to-Image Generation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.757941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.757941Z digest=sha256:a87a88e6758e700ba8c69019922bea4654231dd6e6d20e75c9db35bb3c402bf3

Observation a31a52af-6750-4c01-9158-4dd633caea2b · outbound

This paper cites Evaluating object hallucination in large vision-language models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Evaluating object hallucination in large vision-language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.763518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.763518Z digest=sha256:ab52b588695a6006b164397145b68bb6a64dea7930b54472c00dfd8948662f9d

Observation 214cf737-bda8-40cc-b2f3-d8c40d832ef4 · outbound

This paper cites Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Uni-MoE: Scaling Unified Multimodal LLMs with Mixture of Experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.768526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.768526Z digest=sha256:7da460992c213cff639f1c987634251dd8e49eb763000bbd5197f412808d2c8a

Observation 94cc0d62-b2ba-4968-84f9-fe064100517e · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.779374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.779374Z digest=sha256:6cb5a819ffab24d4b175122e6c67e37c44a8fbc0473481678199c6aeaf02bdc8

Observation 953ab72b-0ed0-4c7f-9633-134915fe8266 · outbound

This paper cites Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Doll ´ar, and C

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.784849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.784849Z digest=sha256:b11f6a1b22bae9e1fab8f29216572eb3096d34c05e2bd0df49bb361ca1ff79cf

Observation 4a3ea85b-a8eb-462c-95be-3027a66e2f37 · outbound

This paper cites MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MMC: Advancing Multimodal Chart Understanding with Large-scale Instruction Tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.789272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.789272Z digest=sha256:795f7284454ae1dfbc8508a2a87f4ebe2f0ff095932b482e5b3a9b2170df3ff4

Observation 2fa7e85b-d979-4fc3-8779-87d05185bff0 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Improved Baselines with Visual Instruction Tuning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.798985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.798985Z digest=sha256:f186a0b1924c191cda361a8648d72561e6d9a4bb884a63300045066a5698a245

Observation 6542c7df-d06f-48c2-ad87-151e4a6aab20 · outbound

This paper cites Visual instruction tuning.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Visual instruction tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.803667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.803667Z digest=sha256:d558f9d066f904e39ca0ceb776da6ade9fc2281f17c67d46a85e06d5c9ea403d

Observation 18066164-3919-4d33-9fb1-6e4b2a0ec29d · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.808284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.808284Z digest=sha256:fe4b0b4251b1bc780dd4b5254fcc1f64cbedd581eaaf195379f236874be86511

Observation 823748f5-ebdf-47d1-9410-ea0dc575633b · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MMBench: Is Your Multi-modal Model an All-around Player?

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.812762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.812762Z digest=sha256:ad740d77e6842a22d27cc2574eb0113d85dafb7e51e2985061ce11ea8810a628

Observation b231a867-28da-4937-a338-28fd809b4f31 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.817350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.817350Z digest=sha256:4467e5cab0bd039071c0851d982fec72325f8eeef1f01635e2234af9465435f3

Observation 0c6fa849-dae6-4939-9a3b-948912f29079 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.822438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.822438Z digest=sha256:bfc199176367ccad31ce1c8ef2332ca1242c4d95e29ce845be55bd979e5eae00

Observation 95ceb76c-ab1b-4a18-b139-e4de7dfd9faf · outbound

This paper cites Learn to explain: Multimodal reason- ing via thought chains for science question answering.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Learn to explain: Multimodal reason- ing via thought chains for science question answering

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.827314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.827314Z digest=sha256:e7b1cdd4eeec3ce64df6420b774423bd0595b82c8c2cfe5a9784a9489028885b

Observation 759827bb-5dd1-4dfb-b87c-9cc116ef2d52 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.832020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.832020Z digest=sha256:7e945299d859cc367fcc14419ffbeb94563b908b0d7a89788eb2d2b36eba6fd2

Observation ff1527d0-8cf4-40b9-89e2-24b1568aa1ac · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.837887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.837887Z digest=sha256:adbeb5efb887dea374747fe56ed45bf2d2b0c3f5ce0061f7280fe7a1684ad681

Observation 7ab74eb3-4106-4e02-9452-c0710906d1a0 · outbound

This paper cites Megalith-huggingface.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Megalith-huggingface

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.843057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.843057Z digest=sha256:816e08f575c12ba46a7e8f6c6755ea8174059bb86d085b50359dd00c04205277

Observation cb06c487-3230-4df5-bc62-809a502e7238 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.848361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.848361Z digest=sha256:82ef16e5f578cbe8259f006b18c0b43ea19856b9f0743a5ed3441836aafbf67b

Observation 05299ded-ce4b-4456-ac0c-314d96aa490c · outbound

This paper cites Infographicvqa.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Infographicvqa

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.852944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.852944Z digest=sha256:2f6d76b5800d25e699e6847fe710a4de28a75817ab93ebf8d58bf6af69100ad7

Observation 4532c57a-e7b4-43fe-865e-cfe7eed7a93a · outbound

This paper cites Plotqa: Reasoning over scientific plots.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Plotqa: Reasoning over scientific plots

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.857747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.857747Z digest=sha256:c23e9e9b9b92c559c892a4ae6f40f63e29736c2c3d08dcc92359079393831052

Observation 47ba5385-a623-4650-af1f-1098c0ff8435 · outbound

This paper cites Hierarchical Attention Encoder Decoder.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hierarchical Attention Encoder Decoder

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-11T17:00:09.706487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.862592Z digest=sha256:f0df88d569ee26730ef1d89a84e96a056d97d738988b427e162d7f7096cd874d

Observation e27deaa8-8d15-40dd-b367-e53e8c1d7eeb · outbound

This paper cites Datamux: Data multiplexing for neu- ral networks.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Datamux: Data multiplexing for neu- ral networks

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.867655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.867655Z digest=sha256:54db1b63f46c8ec3739f5a713367fd8113d80bf4da7b1f40d1fa65d38e54ad9a

Observation 99e1cc0f-5b05-44ea-8e9b-3ed3b673eaa4 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.872326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.872326Z digest=sha256:208b6dd8d891c5c4c2d5897cfbb824a2d3af64caec427fee754b07b846889946

Observation c80ed520-c997-4125-ba85-ebcd734c8541 · outbound

This paper cites GPT-4 Technical Report.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding GPT-4 Technical Report

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.876788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.876788Z digest=sha256:a0920558a343102b966248eedc8e3dd7148e8fd1ed15fae48a2821a619b09952

Observation a3f4e258-fe17-4fc3-9245-f566b3cb439d · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.881574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.881574Z digest=sha256:ea898bc6fa86cbeceb225210820d30bba1a1eb33c2045755d4d13fe72597fe04

Observation 3dbf7f94-e3c6-46e9-809a-30f5351066a6 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.885916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.885916Z digest=sha256:901aa86ad2be07d3a776ba6f44cdf4230401f5e7215ec3d6870a69ab5f331f51

Observation 0fbfb2d1-850a-4ddd-921c-8a66a008e2d0 · outbound

This paper cites Learning transferable visual models from natural language supervision.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Learning transferable visual models from natural language supervision

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.892509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.892509Z digest=sha256:eefbba33e87430be7832aa99fca33753cedc722d6ba101a581a6019dd9118217

Observation 91b37229-ae72-4618-938b-b506c60cb5a2 · outbound

This paper cites Zero-shot text-to-image generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Zero-shot text-to-image generation

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.897880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.897880Z digest=sha256:6c383bed22daa1f83fcbb8113d167b402cf42de9f96055d8a456c5e4c0b9af86

Observation c4a80815-c067-4f9d-b21b-d067927c4c59 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.902780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.902780Z digest=sha256:e7b7361643fd70f933b4b73322275cafa81a6a9b1e0b7c0a588ca031795147ac

Observation 3ef2d716-f786-4666-82af-c461d8c1c00b · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding High-resolution image synthesis with latent diffusion models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.907682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.907682Z digest=sha256:c1b2b28d8bfef15982a996751cc048450bd7ffbb06d556b05e1938eea65cc788

Observation e68d3105-79b6-4e27-b487-bfd73f660c2f · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Photorealistic text-to-image diffusion models with deep language understanding

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.912210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.912210Z digest=sha256:1a01fda83533ee32fb5e1773c75f1e9b6a2400c143999288f75178477467fd66

Observation 2332029b-12ff-4e04-bf90-da2ddc051afe · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.648724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.916986Z digest=sha256:16a4184e36d656fb9450bf743eee4c31da41824503ac177e72e058300d8b303e

Observation c8e5273a-56d5-46f8-a807-040844f3d113 · outbound

This paper cites Laion coco: 600m synthetic captions from laion2b-en.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Laion coco: 600m synthetic captions from laion2b-en

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.629255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.921920Z digest=sha256:30c88df5d51053f71547b183bce8f4d829c9faed7f189973c9330426f0a8a042

Observation 65db8cd2-ac47-4fec-a726-1397959be663 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Objects365: A large-scale, high-quality dataset for object detection

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.609627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.926609Z digest=sha256:fcbc6b1a1f75d6934e372b3a6f75d411151ad9bb8af22a10bb9cc5d8d5ea1ff8

Observation 6cbf65a4-a649-4f50-94c6-09e1b6a3910e · outbound

This paper cites Icdar2017 competition on reading chinese text in the wild (rctw-17).

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar2017 competition on reading chinese text in the wild (rctw-17)

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.931722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.931722Z digest=sha256:1c53b5cf591f08821095e0d7ca50273f44c2017b528d2ca7bb923cadfb6c9154

Observation 959718d6-7bdb-469a-8019-85e987e29bc0 · outbound

This paper cites Textcaps: A dataset for image caption- ing with reading comprehension.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Textcaps: A dataset for image caption- ing with reading comprehension

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.574575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.936911Z digest=sha256:0d07912dc71e6fe59d9915eeb6dc2199c3feb57f39cc336f28e1ab4060bfe0dd

Observation 768ed25d-a7bb-4c4f-a38a-6fbe8c187ec9 · outbound

This paper cites Towards VQA models that can read.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Towards VQA models that can read

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.941854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.941854Z digest=sha256:f8c70d2223a1d5b4b19cc7f2ea499d441b77c1af21c566643d7a94b2ae4a2d04

Observation 1bd3c2e8-7aa1-4505-9165-c7ef25287978 · outbound

This paper cites Textocr: Towards large- scale end-to-end reasoning for arbitrary-shaped scene text.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Textocr: Towards large- scale end-to-end reasoning for arbitrary-shaped scene text

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.946934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.946934Z digest=sha256:9183cdd6ad4be2e1136eb8debe4f90de22c289d19344b96d88a61a147a2b7650

Observation 96cb2e30-59aa-4257-a0c6-f3a388c19d7c · outbound

This paper cites Journeydb: A benchmark for genera- tive image understanding.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Journeydb: A benchmark for genera- tive image understanding

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.528887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.951796Z digest=sha256:0c34531c1332e7edbf53f8af0bc15373e5bb69e29c67677ebc7ab91ac50fe717

Observation 079cfd98-7c10-4185-bd09-ef0be4b338b8 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.956756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.956756Z digest=sha256:b321179bb0ceaa2540746d173a383b0426a80b0a4852492f7f8bc10cc1bf7ba6

Observation 402872a8-4df9-4a19-a347-359994e98ea9 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Generative Multimodal Models are In-Context Learners

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.961656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.961656Z digest=sha256:b752821a168609c83143e94f83c04ca47f2d7bd4d4c6ac6d324be0469a095a68

Observation f18f29e1-cc4d-4038-9989-a633c6a96516 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Emu: Generative Pretraining in Multimodality

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.967050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.967050Z digest=sha256:9fe672c31cc99e169c7b5d5cffd99c24c102ae123ea0f61c75de8c1ad0feffcc

Observation c0a777c7-9113-4297-9fed-e80617cdc55f · outbound

This paper cites Generative pretraining in mul- timodality.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Generative pretraining in mul- timodality

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.507781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.972343Z digest=sha256:d4e8236dc4985ace051e77244c4cb4477b06dc8f2555c2a99614dede16eedbe7

Observation f8b244ba-eaed-4688-b8a0-9e6bee3a2ec6 · outbound

This paper cites Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Icdar 2019 competition on large-scale street view text with partial labeling-rrc-lsvt

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.488918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:08.976643Z digest=sha256:3654bddb4861b64a808c0f27badbb0ee8f5a3665add971e91c44ee30ce0c9e8c

Observation 51119b44-e400-4c8a-a388-6b047d62fb0d · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gemini: A Family of Highly Capable Multimodal Models

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.981015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.981015Z digest=sha256:cd530dcb091a6e98c7268e2211b5ba70197f043772c54e4f2158e9a725287f6a

Observation fe5c4185-b6e7-4225-be54-3b6011fb1680 · outbound

This paper cites MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MM-Interleaved: Interleaved Image-Text Generative Modeling via Multi-modal Feature Synchronizer

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.990779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.990779Z digest=sha256:a42149f7f573f188f597bbbd3803352907ad90fb058ada4c0d57e195139297d8

Observation 4d37afe2-c69f-487d-a723-3f19ffceecc5 · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:08.995294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:08.995294Z digest=sha256:20963743be940654b11f109e92a7c4898dfaa93f7692509433ef91f1482118a3

Observation 07d95d67-17bd-4b36-9dbb-aa0a50a21013 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding LLaMA: Open and Efficient Foundation Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.000066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.000066Z digest=sha256:be2487d90e70bbc81285ac463461d4bc80b76c6940ac3c6058a2dcf04be851ed

Observation 156dff87-32bd-4c4b-8dfa-5b636ff75aab · outbound

This paper cites Unsplash Dataset.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Unsplash Dataset

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.466057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:09.004917Z digest=sha256:90dc8e69fce83337988b37537cf10cafbfc8058a746d686e8828dd352b6792af

Observation b3ce1c53-7158-47fc-8175-0cead03127b1 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.444518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:09.009630Z digest=sha256:473216e392b7295376ca9b36af0adfd5a44d7190413f83b451dd6f8bdf7abcef

Observation 7a9bf386-814f-4c73-8d83-2b4bf614e94c · outbound

This paper cites COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding COCO-Text: Dataset and Benchmark for Text Detection and Recognition in Natural Images

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.014394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.014394Z digest=sha256:455157fa28f0bed637990883b872f078dc9008af55deeaa45c8b15893ec38824

Observation 387791ba-366c-465e-bc13-02b5434d7dd9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.020070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.020070Z digest=sha256:fb3b3ead69a93f80b314423ded3f4e172a6da9089ad3d12e2299cdd77388885d

Observation d0eaa8d3-4c83-47aa-8133-c6cca7ca74b2 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding CogVLM: Visual Expert for Pretrained Language Models

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.024739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.024739Z digest=sha256:61788e555ba0ae9a36c0272f6fd95776254507752f4bc98c7b95ce9fd73bdb9e

Observation 4869e8a8-b6a5-4e97-a538-d665ac4851d4 · outbound

This paper cites The all-seeing project: Towards panoptic visual recognition and understanding of the open world.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding The all-seeing project: Towards panoptic visual recognition and understanding of the open world

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.424785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:09.030050Z digest=sha256:99bd14c9f7644a97cdbe0c0d4ddaa894f763d5fc45f751f8a2f8b14061f6f699

Observation a2ae4707-6c25-409e-9d45-af722ddd40c2 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Emu3: Next-Token Prediction is All You Need

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.035078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.035078Z digest=sha256:b3c62aaf8c21b47df5c2923a45d7bfdc6e135329b9b52aa9b69f79f036fa7751

Observation 65f2ae82-5b65-47ab-aca6-7a247cb6a623 · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.041397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.041397Z digest=sha256:c2e0edca7c877d96fc9da3b7e3e095f6591bab4e9f04e96695674499351c0d6a

Observation 147f3c0c-fba2-47d6-abdf-32112d9451b2 · outbound

This paper cites NExT-GPT: Any-to-Any Multimodal LLM.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding NExT-GPT: Any-to-Any Multimodal LLM

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.047135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.047135Z digest=sha256:d3df4366be6155b341ef6ee63690795c264a2c444a176175933e156d0107956c

Observation 2ae3b315-0142-423a-aba0-ef07ecaded38 · outbound

This paper cites VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding VILA-U: a Unified Foundation Model Integrating Visual Understanding and Generation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.052005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.052005Z digest=sha256:f1ae90af27d7ea2eb400584960bc4a5f2174a2d8647700cabb03f5190b03d0c1

Observation c9b940be-4569-464c-bd13-4b9f31ba7cea · outbound

This paper cites OmniGen: Unified Image Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding OmniGen: Unified Image Generation

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.057044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.057044Z digest=sha256:e7d5ad41eb38f12ef9701cd54bc3c3dad9e5f3d032515da88f76b3cdc661bd8d

Observation ccf93bb2-e6cb-4609-9403-c9770a4acae3 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.062323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.062323Z digest=sha256:6dc8c36f60efba4e9041a67006225dfcfa9f434b204e4b0b1816b87a758cbef4

Observation 0048fa3f-a283-4fc2-a2ce-afe0b736b402 · outbound

This paper cites Raphael: Text-to- image generation via large mixture of diffusion paths.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Raphael: Text-to- image generation via large mixture of diffusion paths

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.396178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:09.067568Z digest=sha256:7009793ea609080f164e86937b4df1817b4dd5af7fd1ad917d79bdd4269c9161

Observation 81d84120-61a7-4c2d-9495-3b0b83f072bf · outbound

This paper cites Qwen2 Technical Report.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Qwen2 Technical Report

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.073175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.073175Z digest=sha256:3933c72c3906077778091439badcbc1c3177914c22d8a41c0ef20403f598117f

Observation 8ab5983e-50a0-43fa-8299-6bbcaed61542 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.078621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.078621Z digest=sha256:5c3d6920b34c82252eee37f8def8e14e284691035ce888dba8193f48b18578e1

Observation 8538b517-b6c4-42c0-9ae9-3b67a9b93a03 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.083555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.083555Z digest=sha256:52d9a2a9161b3e1cc468f359a716a938480fd41da50680b3617736c0434fd56e

Observation a03121b6-dd7a-4630-b6a9-73e3fb28ea33 · outbound

This paper cites Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Scaling Autoregressive Multi-Modal Models: Pretraining and Instruction Tuning

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.088367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.088367Z digest=sha256:8625a5567c81349ab014e3152641b64ed14397cdac313765a1ea47158fcd2138

Observation 9112f0d1-801f-416b-ab8e-2b4b4d0bad98 · outbound

This paper cites Megabyte: Predict- ing million-byte sequences with multiscale transformers.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding Megabyte: Predict- ing million-byte sequences with multiscale transformers

Reference 102

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T17:00:10.378462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T17:00:09.092934Z digest=sha256:0d3155c14bcbcd261ef64557a354f90e6c658db19dec078c7d496e2c1ff87700

Observation 8e052247-65f5-4fff-be4d-7aac0e180649 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-11T17:00:09.097663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:00:09.097663Z digest=sha256:70555481505508ac554b02521861ead9e357d0b166c2830956b1b36ed3d27365

Pith citing papers

Observation c94749cc-882d-4423-99fc-664959a4ca99 · inbound

Liquid: Language Models are Scalable and Unified Multi-modal Generators cites this paper.

Liquid: Language Models are Scalable and Unified Multi-modal Generators SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T21:35:40.295495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:35:40.295495Z digest=sha256:c186271cf617300c40a86f0e296142a1f9c38d20a5187db361b958a85bc331e9

Observation 853fab78-30c7-40f5-af36-9aec3406c865 · inbound

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding cites this paper.

Parameter-Inverted Image Pyramid Networks for Visual Perception and Multimodal Understanding SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:39:35.466176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:39:35.466176Z digest=sha256:d90a98e1fe946e608bb4ff6a8929dbafe107e8acb17573f65635b06244bd351c

Observation 0bec555d-1264-4718-b863-d5dba6560ffc · inbound

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths cites this paper.

UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 324

Resolution
unresolved
no resolver link, observed 2026-08-08T15:24:46.358407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T15:24:46.358407Z digest=sha256:216f25c0a5f4d8d7bf48b5e9f562366309d35ec486c13c7518f0735caa04e6d4

Observation d7bfb031-8595-4721-b910-a1f59cedd831 · inbound

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation cites this paper.

WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:24:27.644062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T16:24:27.407376Z digest=sha256:3f690cf982b14e4ed09b32962d94e48cc8297baa0b39a8198d010f53a35a4f33

Observation d4fa1d2f-af8a-4a60-aea4-08ba69542773 · inbound

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning cites this paper.

UniRL: Self-Improving Unified Multimodal Models via Supervised and Reinforcement Learning SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T12:53:23.378411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:53:23.378411Z digest=sha256:502934b9a3b3348e76403060b225610e40bcc01b894ff1b2ca7fbb13fdb577c8

Observation e06dda0f-75f6-4bd1-8c8d-770a35ed0620 · inbound

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation cites this paper.

OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:44:14.957297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:44:14.957297Z digest=sha256:40f77ce3747d0ae8cb4d0ac82dd1d18dd3d643191bec1c5cff25ba2d3725b7cb

Observation 5c8e9261-e7f5-452b-b842-2ef5ab478aef · inbound

Show-o2: Improved Native Unified Multimodal Models cites this paper.

Show-o2: Improved Native Unified Multimodal Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-12T18:51:15.918479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T18:51:15.428692Z digest=sha256:c12ef0e6c832f98fab8b51d81ece3c4ae974fb8fc01294c6207b06fd99181bbf

Observation 718b4b77-42d6-4f16-baa8-b702afb4b3c8 · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T16:49:57.751825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:49:57.751825Z digest=sha256:11101aacfc4b2ad0123bbfaeb95d4370332690ef878892e168b1510f843805c0

Observation 1547e029-5114-4d1b-8b0e-5954bf0be3ca · inbound

Discrimination Is Generation: Unifying Ranking and Retrieval from a Tokenizer Perspective cites this paper.

Discrimination Is Generation: Unifying Ranking and Retrieval from a Tokenizer Perspective SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token Folding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:50:24.283929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:47:36.536059Z digest=sha256:ff3ce3b5feb41fbc537ff19479fbc8aa21acb5bc08ff0c320aa978fd1f27cc8a