Pith. sign in

Paper Citation Record · LEDGER

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 3 inbound Pith citation observations for arXiv:2605.12309.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.12309 v2

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-19T16:53:11.708868Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T17:23:16.226474Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

44 of 44 outbound references displayed

  • verified exact18
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8925a8f0-1eac-4acc-9206-b0aa14c0c37d · outbound

This paper cites GPT-4 Technical Report.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models GPT-4 Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.355927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:da9fcb79415ef1cce40af376616bd478cdb310c4df946cf5348cf71388d38d1d

Observation 57835968-8203-488b-85b5-8fb4a9e81436 · outbound

This paper cites Improving image generation with better captions.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Improving image generation with better captions

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.063119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:3f76d25193f22748d47ea0ea8fd44fc0224b3d99d3d6c894e2d121f12c7bfa03

Observation 93c6eadc-433d-43c1-8739-fb5db997e5ad · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.067222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:2a1192130ff06b31559216ec389b04a4fe215dc4bb56cb1ad256c8e6444d3122

Observation 2fb7338c-1f67-4d00-93f5-a5513ab3418a · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.361439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:2adec8db8816692fde0f92d2434cf2a4d772be66a4d0c4b7502b24cf106b7489

Observation 409de80c-f4d7-43a7-a20a-96dd8d20815d · outbound

This paper cites Diffusion models in vision: A survey.TPAMI.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Diffusion models in vision: A survey.TPAMI

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.065027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:78844dc22159be6f0ebcc3d05420c720b6f075beee7929dfe97338c0eec972f3

Observation acf3a854-831c-4530-ad46-3314c1458c8f · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.069150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c3ab81f4032dddde76595e37e79fc4e3f6edbf7adf5b5e782c0fe4f55f49f211

Observation 2edc038a-cb93-411e-b432-42c1babb6af3 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emerging Properties in Unified Multimodal Pretraining

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.358740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:939942e40c5af54bbd20e7f929e0bc2650c538965235b1bce9ad7ac98fdc3da9

Observation fb4f0b5d-4128-4980-9b24-fdb4e39065a6 · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.091236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:15a2a633a8f17d310cd0fb60ca9179175847e7eb0a52d86c3b2f290a711df2b0

Observation 92304e4c-3f92-4a58-89d6-eadbfdbb90d2 · outbound

This paper cites Gemini 3 pro image (nano banana pro).

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini 3 pro image (nano banana pro)

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.089363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:458e05b68814fe4566f66b21c15bdc55b2611b562162c3424c512cd092edb97f

Observation 512906cd-a925-4aaa-a898-578662eee7c6 · outbound

This paper cites Understanding and harnessing sparsity in unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Understanding and harnessing sparsity in unified multimodal models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.396959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e74346e4d7d2edf4efd3d093f5d605afd487627ef6a841b7abd4ea08bb6829e9

Observation f7194ef7-c9ca-4ffb-a6d5-292cffebf5ef · outbound

This paper cites Flux.https://github.com/black-forest-labs/flux.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Flux.https://github.com/black-forest-labs/flux

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.095103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:bea391360c00339ab199bb832e187ffa135bdeaf7609d9123c69ff650e9118c5

Observation 1f75a905-337c-41eb-b527-ae6874f6692a · outbound

This paper cites PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models PlanViz: Evaluating Planning-Oriented Image Generation and Editing for Computer-Use Tasks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.367552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:a7bc3a754b61322fe7c9149abd03419b4543b6ad579e28240934f715d2713335

Observation a51cdea8-1a5b-4a2b-a019-274e5fd45dca · outbound

This paper cites Dual diffusion for unified image generation and understanding.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Dual diffusion for unified image generation and understanding

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.105285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:a3812ab122ce5e4d34450e0a30195c3cb57629cdce34835f0f64b93ed4828362

Observation b7f3a2a2-d2e0-43da-9e26-f4b372086679 · outbound

This paper cites Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rover: Benchmarking reciprocal cross- modal reasoning for omnimodal generation

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.400396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:92f038118f1689c29938617ff8db94c7e0d647aab910eb951fcb91d0c1a1de78

Observation c2acda8f-bb3d-4509-bc3f-7f23171c38f5 · outbound

This paper cites Visual instruction tuning.NeurIPS.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Visual instruction tuning.NeurIPS

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.097712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:86967c71473955ff75af85a7d52e80cf7af63e98c20c037e3aab7a2128e7116a

Observation 702ed9fb-c920-4ae1-8f1a-07f44c4fd09d · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Step1X-Edit: A Practical Framework for General Image Editing

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.379475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:1b447a004ee1e909758c6191e011354fdc4ad1fe09de0b0ae0457ab95f2f8c66

Observation 396631f7-befa-4636-88a1-800ac4718700 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InECCV.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Mmbench: Is your multi-modal model an all-around player? InECCV

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.109278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:547e18171337477997da3c20a897609c7ac4556bd810f51a34d262636e5d8695

Observation 26e843a3-6b1e-45bc-9c71-5c2a53ef71b9 · outbound

This paper cites UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models UniMoD: Efficient Unified Multimodal Transformers with Mixture-of-Depths

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.364921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:e08e54f2715a58b7fe9ec823333704fc3bce0443a82412edcbd4e199bd1d3878

Observation b6f3980e-1133-4e7e-b1c0-a35a00f4643e · outbound

This paper cites Introducing our latest image generation model in the api.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Introducing our latest image generation model in the api

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.081011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:14547e47eaee5ed0fb7439b5476ac149ae176e01f81a583de69664134c82d47e

Observation c7dafb59-dbd9-49c7-8c81-e274df65986c · outbound

This paper cites Wiseedit: Benchmarking cognition-and creativity-informed image editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Wiseedit: Benchmarking cognition-and creativity-informed image editing

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.370572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:27c79a2a576828a4c93bc315225c605d8daba9a794da4e66bb051aa7e2a40466

Observation c257be5a-160e-4068-9868-a496e09957f2 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models High-resolution image synthesis with latent diffusion models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.073244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:70f42d40f1422db583b5ebec74eb6dad878cfc166b2a497e1aeb116d1e3b8c9b

Observation d9698088-0486-4840-9bb6-dfd79191744f · outbound

This paper cites Holitom: Holistic token merging for fast video large language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Holitom: Holistic token merging for fast video large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.075719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:6927ad9c8b2fc6042703529b0ab97181735effb422bf66428a80b15380fb203c

Observation 5279849e-2857-4484-85a1-a1549090b85c · outbound

This paper cites A survey of token compression for efficient multimodal large language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models A survey of token compression for efficient multimodal large language models.TMLR

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.071140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c7b357429e56560786f08fbb1de053fbd9302b8612e4d34703fb04b2d80aa3a8

Observation 6648d7ce-bd9b-43ba-a4a2-2f24da9d50cb · outbound

This paper cites Less is more: A simple yet effective token reduction method for efficient multi-modal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Less is more: A simple yet effective token reduction method for efficient multi-modal llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.079133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:1825fe52ae0a6667424c10d9e50b353f3b9480ae00151c06d8375bd3f2ce8df1

Observation 65fe70a0-d0bf-4b63-8b6e-dc285978a3ab · outbound

This paper cites Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Ivc- prune: Revealing the implicit visual coordinates in lvlms for vision token pruning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.376802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c73db73a8c3fcb2efa1a92691cd70d648bb4fe7880e62ac49e2df06c1c0b9a25

Observation 90089d6c-0c9a-4f69-a0a6-6688a39f8eb4 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.394111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:7b0aa9ec1537d9947cc1d91550f08c81621468c320338816451f9d6f26a6fc91

Observation e8fc5633-70f0-465c-9a87-538748b6a09c · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Gemini: A Family of Highly Capable Multimodal Models

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.391340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c15465aba2552bbe00253ed8cb3ed4b4f85182f4dc7bf625b29da4aa6f78fdb0

Observation f4b54679-5e4b-423c-88e0-ec99631350fe · outbound

This paper cites Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Internvl-u: Democratizing unified multimodal models for understanding, reasoning, generation and editing

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.385307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:33dd639699f41599cd513296305aa5d511962e21297d62ceb61df9e4fb223f40

Observation 8b167d6f-4d5c-46f9-9db4-6bdd10dedf49 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.085232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:0f634ba6367bc461e3a1c38768463c9d3627bfe6bbdccdbcce8669413529570a

Observation ce83964d-a53f-4935-bb8e-7df1de413c51 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.382343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:c881367cad46cb27a9643e6b8649a454d3d6716a3b691992719d95620cf89c24

Observation 4b4c5688-3ebe-4e3c-a403-ffca9069dd12 · outbound

This paper cites RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.403042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f08f848b891e58006fcc7ad19d5d66e00b12f3e20afbf7eb7708f52ad381addc

Observation 9dfbf07c-325d-4a9d-88ea-608045d23a5c · outbound

This paper cites Emergent hierarchical reasoning in llms through reinforcement learning.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Emergent hierarchical reasoning in llms through reinforcement learning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:57:40.406050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:28060053fe69c02898c57b731b92c7ea56cbc54e82862e818183aabd805550ba

Observation 263065b5-48c9-4c98-83be-91bec573a427 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.388168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:21abca776b88d38a62dbeb435fe01831c26953cfa0f79a07d9d092e925369ca5

Observation 1b789c3a-4809-4c74-8ac0-7956950501e0 · outbound

This paper cites Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Token pruning in multimodal large language models: Are we solving the right problem? InACL Findings

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.107312Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f588b835f81fc3898c7e08dca5a64982d3a6346886f66530f44693311d31bf57

Observation 8bbc5f82-5fe6-497d-9c08-99d64483cfda · outbound

This paper cites Janus: Decoupling visual encoding for unified multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Janus: Decoupling visual encoding for unified multimodal understanding and generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.111113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:4b3c23aac41b6a2f733aa88df2f531baf1ac2f03c76fcdfffe56ed199e810820

Observation 68a2200a-c987-48c0-ad26-275c56e6a2bc · outbound

This paper cites Kris-bench: Benchmarking next-level intelligent image editing models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Kris-bench: Benchmarking next-level intelligent image editing models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.103247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:af75265b1715c490970c98446be0359eed6e4a9e6a340ce7ec391797d4d964f8

Observation 06362bac-ed3e-497d-be0b-3e558285b072 · outbound

This paper cites Announcing grok-1.5.https://x.ai/news/grok-1.5.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Announcing grok-1.5.https://x.ai/news/grok-1.5

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.061118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:667dade9c293ba770684322225ab3fcadeddb7919f2c6c233d36d93a6b6772fd

Observation 0e9686c9-6da3-40ea-950f-e098201c1668 · outbound

This paper cites Show-o: One single transformer to unify multimodal understanding and generation.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o: One single transformer to unify multimodal understanding and generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.093044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:ea3a09f1798764feb02c6a082769ea4ee5d5eb06fa1b096c5bcabd61844f8f1e

Observation 653e866e-61f3-45bd-8d94-dff9c366794f · outbound

This paper cites Show-o2: Improved native unified multimodal models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Show-o2: Improved native unified multimodal models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.099570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:5fbbb69eca38ecf71ba54dee645dc7b2f2277b55753d04502f9e34cc29176bb8

Observation ddfbf650-9dd1-45aa-a626-43c53268e365 · outbound

This paper cites Conical visual concentration for efficient large vision-language models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Conical visual concentration for efficient large vision-language models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.101372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:f7e87b2bb38358a0b3e9a4697a4d76fd1f7886afd14c7c71cda2e2321a0a28d3

Observation 369b972d-879d-4699-b2f2-69434741e0c2 · outbound

This paper cites Rethinking visual token reduction in lvlms under cross-modal misalignment.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Rethinking visual token reduction in lvlms under cross-modal misalignment

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.112952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:b456f3bf0b8db2e530b73e2ec38c5b661377659999d3eecc05f030ee629d8663

Observation 333c58d3-8dac-461e-8621-c8f93a8cf2db · outbound

This paper cites Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Vscan: Rethinking visual token reduction for efficient large vision-language models.TMLR

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.083000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:55a96e989b9137355ca51df271097006cd0b5e340e327929d2134abcf0d86e00

Observation c0103691-5230-440f-a923-24eebb1f3db9 · outbound

This paper cites Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models Envisioning beyond the pixels: Bench- marking reasoning-informed visual editing

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-19T16:57:41.087497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:791f610fb949ab70fa42b08be209ca6a834caabfcd0cbdb565efa199283c49fd

Observation f49e966b-04a2-403c-9ca8-e08a3a2ac678 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:57:40.373501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T16:53:11.708868Z digest=sha256:af962914f9dc89617cf0077412f97614b9b370fcec7cd38d2e392f8f7737c4f1

Pith citing papers

Observation 66fef0c4-b662-4749-8b0b-bb26e25f41b4 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T19:41:47.727930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T19:41:47.727930Z digest=sha256:efe163e5c9269c5ea3104c1a21489903cc1b3857dfbce7a1ef1a75c239b92d48

Observation cb1dcba1-9b29-49f2-8e36-bda0ef0f3212 · inbound

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing cites this paper.

Cross-Branch Conflict as a Shield: Safeguarding Facial Identities in Unified Multimodal Image Editing G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T04:17:41.520667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T04:17:41.520667Z digest=sha256:a20ed6e4a2a661c2b581fe8854f44c004736b0f5e2a520b9cc591fb2690d71ec

Observation f6b5f369-d12e-4297-b3e7-68f9c8070272 · inbound

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs cites this paper.

ET-Prune: Evidence-Aware Dynamic Budgeting for Visual Token Pruning in Text-Rich MLLMs G$^2$TR: Generation-Guided Visual Token Reduction for Separate-Encoder Unified Multimodal Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T17:23:16.226474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:23:16.226474Z digest=sha256:578d21990d1fd82cfbb1cec9efc9f7a7a67e47c1a858f3112b360a0d998c28b2