Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T16:53:11.708868Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.12309.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-19T16:53:11.708868Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T17:23:16.226474Z
A source-named dated measurement, never combined with another source.
Source: cited_works
44 of 44 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8925a8f0-1eac-4acc-9206-b0aa14c0c37d · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 57835968-8203-488b-85b5-8fb4a9e81436 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Improving image generation with better captions
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93c6eadc-433d-43c1-8739-fb5db997e5ad · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2fb7338c-1f67-4d00-93f5-a5513ab3418a · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 409de80c-f4d7-43a7-a20a-96dd8d20815d · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Diffusion models in vision: A survey.TPAMI
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation acf3a854-831c-4530-ad46-3314c1458c8f · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2edc038a-cb93-411e-b432-42c1babb6af3 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb4f0b5d-4128-4980-9b24-fdb4e39065a6 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 92304e4c-3f92-4a58-89d6-eadbfdbb90d2 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini 3 pro image (nano banana pro)
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 512906cd-a925-4aaa-a898-578662eee7c6 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Understanding and harnessing sparsity in unified multimodal models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f7194ef7-c9ca-4ffb-a6d5-292cffebf5ef · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flux.https://github.com/black-forest-labs/flux
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1f75a905-337c-41eb-b527-ae6874f6692a · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a51cdea8-1a5b-4a2b-a019-274e5fd45dca · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Dual diffusion for unified image generation and understanding
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7f3a2a2-d2e0-43da-9e26-f4b372086679 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c2acda8f-bb3d-4509-bc3f-7f23171c38f5 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Visual instruction tuning.NeurIPS
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 702ed9fb-c920-4ae1-8f1a-07f44c4fd09d · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Step1X-Edit: A Practical Framework for General Image Editing
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 396631f7-befa-4636-88a1-800ac4718700 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mmbench: Is your multi-modal model an all-around player? InECCV
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b6f3980e-1133-4e7e-b1c0-a35a00f4643e · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Introducing our latest image generation model in the api
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c7dafb59-dbd9-49c7-8c81-e274df65986c · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Wiseedit: Benchmarking cognition-and creativity-informed image editing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c257be5a-160e-4068-9868-a496e09957f2 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models High-resolution image synthesis with latent diffusion models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9698088-0486-4840-9bb6-dfd79191744f · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Holitom: Holistic token merging for fast video large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5279849e-2857-4484-85a1-a1549090b85c · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models A survey of token compression for efficient multimodal large language models.TMLR
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6648d7ce-bd9b-43ba-a4a2-2f24da9d50cb · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Less is more: A simple yet effective token reduction method for efficient multi-modal llms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 65fe70a0-d0bf-4b63-8b6e-dc285978a3ab · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 90089d6c-0c9a-4f69-a0a6-6688a39f8eb4 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e8fc5633-70f0-465c-9a87-538748b6a09c · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini: A Family of Highly Capable Multimodal Models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f4b54679-5e4b-423c-88e0-ec99631350fe · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8b167d6f-4d5c-46f9-9db4-6bdd10dedf49 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ce83964d-a53f-4935-bb8e-7df1de413c51 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4b4c5688-3ebe-4e3c-a403-ffca9069dd12 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9dfbf07c-325d-4a9d-88ea-608045d23a5c · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emergent hierarchical reasoning in llms through reinforcement learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 263065b5-48c9-4c98-83be-91bec573a427 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1b789c3a-4809-4c74-8ac0-7956950501e0 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8bbc5f82-5fe6-497d-9c08-99d64483cfda · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 68a2200a-c987-48c0-ad26-275c56e6a2bc · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Kris-bench: Benchmarking next-level intelligent image editing models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 06362bac-ed3e-497d-be0b-3e558285b072 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Announcing grok-1.5.https://x.ai/news/grok-1.5
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e9686c9-6da3-40ea-950f-e098201c1668 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 653e866e-61f3-45bd-8d94-dff9c366794f · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o2: Improved native unified multimodal models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ddfbf650-9dd1-45aa-a626-43c53268e365 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Conical visual concentration for efficient large vision-language models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 369b972d-879d-4699-b2f2-69434741e0c2 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rethinking visual token reduction in lvlms under cross-modal misalignment
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 333c58d3-8dac-461e-8621-c8f93a8cf2db · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c0103691-5230-440f-a923-24eebb1f3db9 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f49e966b-04a2-403c-9ca8-e08a3a2ac678 · outbound
G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 66fef0c4-b662-4749-8b0b-bb26e25f41b4 · inbound
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb1dcba1-9b29-49f2-8e36-bda0ef0f3212 · inbound
Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6b5f369-d12e-4297-b3e7-68f9c8070272 · inbound
ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.