Pith. sign in

Paper Citation Record · LEDGER

See What You Are Told: Visual Attention Sink in Large Multimodal Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2503.03321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.03321 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:28.622401Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:31.045973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26f19368-69be-4297-ad00-dc9f67a84f3a · inbound

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations cites this paper.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.622401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.622401Z digest=sha256:c37c3dd76878bf67271eb79a9ed7e85cb4fa7de21cb694c1a04868ea4b9d9ab6

Observation e6765169-2bf6-441f-932a-d0bc6ad9374f · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:23.816759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:23.816759Z digest=sha256:3366ac9abee5a93611d8f4dfd56306384c9f3e0c9c5f3dfacb43f26c3aff5ad9

Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · inbound

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features cites this paper.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.290792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.290792Z digest=sha256:c48d9b6432c7cb0c72652b056f8e333b4d54415b9c7073dce16914c08af02c6e

Observation faa9854e-6bc5-4828-aed1-bb21b25fe44d · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:14:22.812861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T21:11:47.804851Z digest=sha256:1c54b67e5b1dafd69732691c189130a099b2fbd7aa909d15ad98e4334c583e64

Observation 6e7dfff8-e5c2-4649-a669-34d764219291 · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:57.009415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:57.009415Z digest=sha256:c322e2956f44248d7b3aa5040ba913619a19afb9b60013c06c9bed54c93363f6

Observation 3102a7f9-dd18-4a0b-b788-6034d468655c · inbound

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment cites this paper.

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T09:47:23.806393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:47:23.806393Z digest=sha256:3aa16cc0aec7a7a50b340b4822915b08f4e6048d085598bd5e15250e1b9c217c

Observation 863d9f65-ccbb-4e1e-9ae2-fd0b085c9787 · inbound

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation cites this paper.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.218482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.218482Z digest=sha256:8505713ff821072e634bb328f63d2b99a50520ca28797b03c35f739e5c708731

Observation 465b55ba-ba7c-4563-a593-04e432d62705 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.254402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:63522b7281fcbd36220ff7b732d4a9cddbcfa475b9f0eb2dad6bc0ce3689141d

Observation d1d25e1a-0373-4e3b-a027-a3115dc80d91 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T21:44:00.054831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:44:00.054831Z digest=sha256:340dc6e2fe1c945ef083064b0fa6fe6ef920a5b62c58cfa09ef6fbcaf450e572

Observation 36680381-e702-4798-bef3-21fc87835478 · inbound

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions cites this paper.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.265815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:c13274bf85715dbf8405085177aedb19cbdb5a584abe62d04483f8e82a2b26cd

Observation 837504ba-23fd-4c0a-827d-e9080d9327e1 · inbound

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models cites this paper.

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:06:38.105307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T21:05:11.117495Z digest=sha256:c70ea96c4f4868cd2464bbd1d2338cc6429f1f6c130655b40af3c190d95a983d

Observation 019e877f-d788-46fd-a323-8e61986d8f3c · inbound

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models cites this paper.

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:00:04.294537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T11:56:29.743281Z digest=sha256:d878285685ce823b529cee817ddcb123ecde3ad2b310af9ab3d5b57ebec53e29

Observation 34e3fb74-d1dc-4b4e-9f29-d0e0bfefbb2d · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.269739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:0156803acb8c992efe2a0ec1a8fe8aadda92ce1f706a5d6cebe104b2b83d6b78

Observation f38698a3-d098-4a92-b8b8-cbf920c25201 · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.794760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:63c186cea4f22af74a454a255d253a9c957ea68686bae98fe1c99b65f87bb5f3

Observation 2c89199c-f754-48a0-8a7c-b8b831309cbb · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.397647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:3f5223f01a371ecef78de17e01bc9b16aac36104179ac3845a12dcbf99de4aef

Observation 96e39971-fcc6-44f3-9e74-9d7d8f9df146 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.685247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:52c784de238ff4b7e178426ede2ca0ad1cfc566f21a9b933c526f5d2fbcc8b90

Observation 15b448f1-d203-4a4e-871d-8f668ec1eb7d · inbound

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval cites this paper.

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:13.605771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T16:56:52.714346Z digest=sha256:c3270105e8e10a7949bb7c831bd9feed354dfc3740e759756b1743a0a330b02b

Observation 4566a0c8-e827-4395-bf5d-cb8ec66cd7a8 · inbound

Large Vision-Language Models Get Lost in Attention cites this paper.

Large Vision-Language Models Get Lost in Attention See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:10.066217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T11:54:01.224588Z digest=sha256:7358ef6550d86b12d73f17217a52b814c7a21185a38a62a15a05a0ae32be4755

Observation a3d5c165-df40-40d4-9ac9-b47f11a0ca60 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.468869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T11:10:58.225078Z digest=sha256:c034d399fb4c0bf9da75b3c372e6e3cf2343259161cd04697c1431b4c9ec7bd7

Observation 8f45ea5d-edad-4a0f-9064-b4ea98aaa6fe · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.355460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:36:56.248838Z digest=sha256:70b0e7c3fcad22ab34132393d49d9f5c75e146580bcc9e0893e5077ce81fbafc

Observation 145b7c44-0cda-4d95-90c0-ff50baaa55d8 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.456705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:023cc119b6dc13bd69f524fa470e3e2cf5fb4e93fa91543ea1b7343ed06efc0a

Observation 941f5569-b466-421a-8c5b-ffb153b9be89 · inbound

Inference Time Optimization with Confidence Dynamics cites this paper.

Inference Time Optimization with Confidence Dynamics See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.619616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T11:12:46.757296Z digest=sha256:3f107c057a3685075bcbd195188542313f8393ac884dd68024c84d41b9674f28

Observation f7cf9279-2f32-40c9-be26-fe63dbc86178 · inbound

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.800634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:28:35.155099Z digest=sha256:33be393606e052fd36988decb620da0606a1bdf98baf7c1ad6e3124d5000641c

Observation a41cd955-68d3-41d9-a696-11c4074591a9 · inbound

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models cites this paper.

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.409378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T11:27:09.779776Z digest=sha256:92ec9efc717c9e5561da0e78a679088bdec1eec232ad4d4ea7577e68f9e3c1c2

Observation b37d17ce-7aef-4aaa-8d64-58ba4f066f4e · inbound

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads cites this paper.

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.644022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T01:54:13.927736Z digest=sha256:4c3dd750c8ffd308746c7cb4d4d7877a619b8cb6857f92fac6a41adfd0cca783

Observation b039fe84-d7ec-463f-a462-d74ffb2c4031 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.634436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:f237cf5fbad884206aed9bf325002f5726dee353956e6a939c12097115cebfb8

Observation afb174be-9df7-440c-af56-d881b879226d · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.316937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:979815cc33afd6e9119132af2e225eb926eb18d40908487cd47e14497538010c

Observation f663055d-df4d-4e1e-b486-d9cf7ee7eb4a · inbound

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models cites this paper.

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.108386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:35:24.118536Z digest=sha256:97461cfb30c60ba39fd177754523a37bfa187ee940544d01f02ef037d21d5955

Observation d90ea871-600b-40cf-900d-1c1f81e05318 · inbound

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression cites this paper.

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:17:49.042253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T10:22:29.131570Z digest=sha256:8fb7ee8650c73f1e026466619205c4637e50f4728025df0ea893e91c34f2f775

Observation dbfcec07-4c4c-43c2-92b9-4538d13d9812 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.048584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:f06e2eac7275302add051855cccb2937ff0027fd17b7121c12b72f51adbe691e

Observation 5082a377-8813-4fe3-ac05-7b19dd3bb3f9 · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.286659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:88feaa554cdbfbd6c0c016446a0fea8b65815833502d868a2e677c8285fe0adb

Observation 9d60cd50-a443-4612-955b-80f85d85c4cc · inbound

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs cites this paper.

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.486668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:51:01.371390Z digest=sha256:18c77f00619a082e2e19dda5b55e515e0010763996cb8948de5b6ff0eca966f7

Observation 669815dd-0801-42ec-8b36-de554e29f9db · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.247821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:87780de03bd34878478594b0214e5e5fb3e506af2d35b7bd1ebe009298878581

Observation 6ff1c143-7362-4272-bccc-416129289b67 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T05:37:19.632869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T05:37:19.632869Z digest=sha256:7131b3501c149a5401004686300efb124da3d54c4ef6801d360c82012c6e9b1a

Observation e2cba484-7271-45db-ada6-1c7465442908 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:19.935203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:01:19.935203Z digest=sha256:a59a9eff3fa85c1c9de5666e0b7f99bfcd7a5f57203e61991f77c8c3bfacceda

Observation dfaf8944-d82a-4aba-bcd1-997dde3b7c7d · inbound

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings cites this paper.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.730157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.730157Z digest=sha256:82c0c6b38bbf836dba50e473453498b49a998c43f852e9602018f4b24ac0506a

Observation a9d423ba-0d22-4dc5-b4a2-33fed6b76e9e · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:35.906973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:35.906973Z digest=sha256:426ea0080ff78ef9a3b9368bb6fcc4f48b4165e1a495b3eb446f85c82d7fef6e

Observation b5b66a19-8b7a-43b8-b13f-0909d0cc1a4a · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:54.711000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:54.711000Z digest=sha256:a647b4fcbc562a7328c0cf28666e606f05a33ee08b92dd3cd06a1f4d170d3c9e

Observation cd3f7e0f-08b8-4be6-bc56-4994d17e9b8b · inbound

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles cites this paper.

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:37.236788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:37.236788Z digest=sha256:ff84cea6828460f2aa2ed583e6a9068afedb1300b90062592248c57272c39965