Pith. sign in

Paper Citation Record · LEDGER

See What You Are Told: Visual Attention Sink in Large Multimodal Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2503.03321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2503.03321 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:37:58.965659Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:19:31.045973Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 163df10a-8836-4206-ae1b-7df13092ee82 · inbound

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models cites this paper.

Visual Attention Never Fades: Selective Progressive Attention ReCalibration for Detailed Image Captioning in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T15:26:18.833482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T15:26:18.833482Z digest=sha256:ccb49bf618dcc2e92f50b370796f0ed3956b41c9a098501484b999c9574fd938

Observation 26f19368-69be-4297-ad00-dc9f67a84f3a · inbound

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations cites this paper.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:45:28.622401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:28.622401Z digest=sha256:0b46e306fd7d696623182639ffda256b10d839811d03c9600ef04266f0a34beb

Observation e6765169-2bf6-441f-932a-d0bc6ad9374f · inbound

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation cites this paper.

Not All Tokens and Heads Are Equally Important: Dual-Level Attention Intervention for Hallucination Mitigation See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T00:53:23.816759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:53:23.816759Z digest=sha256:c491182d4d5b4c5ea43d7d4df01f3f543f78225a60254681899a2072636d8f3b

Observation e4086921-1ee2-4750-a572-0c1d470f44b9 · inbound

PEVLM: Parallel Encoding for Vision-Language Models cites this paper.

PEVLM: Parallel Encoding for Vision-Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:37:58.965659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:37:58.965659Z digest=sha256:fc56bf175ca1ee60f2bfed3911b7a7ad745e5430c0d582dd323a85e86a9fc130

Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · inbound

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features cites this paper.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.290792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.290792Z digest=sha256:0cf096185e0c2461bcc4539f166ceb2dce9e95415eda7f4c7515e245e661eff6

Observation faa9854e-6bc5-4828-aed1-bb21b25fe44d · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:14:22.812861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T21:11:47.804851Z digest=sha256:84cad0278174dbc159ff3f760e04c4afca3f722e1a10fcb075dc3a96bf8ff82c

Observation 6e7dfff8-e5c2-4649-a669-34d764219291 · inbound

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling cites this paper.

HiDe: Rethinking The Zoom-IN method in High Resolution MLLMs via Hierarchical Decoupling See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T14:42:57.009415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T14:42:57.009415Z digest=sha256:59a97c4a2d6bdaf7a1a61442a01260c4598347e21c673507048b190070a96c03

Observation 3102a7f9-dd18-4a0b-b788-6034d468655c · inbound

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment cites this paper.

Attention Misses Visual Risk: Risk-Adaptive Steering for Multimodal Safety Alignment See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T09:47:23.806393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:47:23.806393Z digest=sha256:ff39e87f9fd4353d994f0c3515f347621bc3e70c90634ece835078c7794f9af5

Observation 863d9f65-ccbb-4e1e-9ae2-fd0b085c9787 · inbound

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation cites this paper.

Capturing Gaze Shifts for Guidance: Cross-Modal Fusion Enhancement for VLM Hallucination Mitigation See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T08:14:36.218482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T08:14:36.218482Z digest=sha256:ae83a5381a137dc17772c288ae84d2ee96fbf47b1e077ce5d401e1d8e593fbc1

Observation 465b55ba-ba7c-4563-a593-04e432d62705 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:54:20.254402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T19:51:04.983299Z digest=sha256:929a6044be588f4cd5f127435b51eed880fad6112b8d00384ec7a49e51910c5a

Observation d1d25e1a-0373-4e3b-a027-a3115dc80d91 · inbound

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs cites this paper.

MVI-Bench: A Comprehensive Benchmark for Evaluating Robustness to Misleading Visual Inputs in LVLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T21:44:00.054831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:44:00.054831Z digest=sha256:f47ebaf01ce5e4d2add1ae8ee73d036b9a08a9ce8d085c312b0176336ad466d8

Observation 36680381-e702-4798-bef3-21fc87835478 · inbound

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions cites this paper.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.265815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:0754f5f13c6a79985885b91612795d0ba6a7e60f2a536752b8474efbce4b360a

Observation 837504ba-23fd-4c0a-827d-e9080d9327e1 · inbound

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models cites this paper.

EAGLE: Expert-Augmented Attention Guidance for Tuning-Free Industrial Anomaly Detection in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:06:38.105307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T21:05:11.117495Z digest=sha256:f3e3289487352547163e92c802095ce15158bec949ec01b2054d3b17ac645a83

Observation 019e877f-d788-46fd-a323-8e61986d8f3c · inbound

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models cites this paper.

Deeper Thought, Weaker Aim: Understanding and Mitigating Perceptual Impairment during Reasoning in Multimodal Large Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:00:04.294537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T11:56:29.743281Z digest=sha256:e7f28b8d9d31fed38a6d52429ace0f6039f00537bea81018228d4deecc8ec35b

Observation 34e3fb74-d1dc-4b4e-9f29-d0e0bfefbb2d · inbound

Counting to Four is still a Chore for VLMs cites this paper.

Counting to Four is still a Chore for VLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:06:03.269739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:39:06.327888Z digest=sha256:467fde2cd824655626506abf3a94fafc70c44fca1a596874c63b7a857dcfc930

Observation f38698a3-d098-4a92-b8b8-cbf920c25201 · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:35:26.794760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:00b63a43771ecfe57132f33542032cabfad6e27a68f2f9520df8e593cb6cba19

Observation 2c89199c-f754-48a0-8a7c-b8b831309cbb · inbound

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs cites this paper.

Sink-Token-Aware Pruning for Fine-Grained Video Understanding in Efficient Video LLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:44:48.397647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:43:44.921189Z digest=sha256:2097618455282ffd1fbe148f9341d7d9404fec1d67ab53cea72d0948f3712996

Observation 96e39971-fcc6-44f3-9e74-9d7d8f9df146 · inbound

Latent Denoising Improves Visual Alignment in Large Multimodal Models cites this paper.

Latent Denoising Improves Visual Alignment in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:09:26.685247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-09T23:07:54.806529Z digest=sha256:a24e627bfffe71540e488e4f917fd29d50391c91792f43b2a6797207e066d68f

Observation 15b448f1-d203-4a4e-871d-8f668ec1eb7d · inbound

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval cites this paper.

Combating Visual Neglect and Semantic Drift in Large Multimodal Models for Enhanced Cross-Modal Retrieval See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:31:13.605771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T16:56:52.714346Z digest=sha256:0af11f64b9e8e4ee8c383c2620c1a1f2c9b42fac831cae05afd6e8663aee3b20

Observation 4566a0c8-e827-4395-bf5d-cb8ec66cd7a8 · inbound

Large Vision-Language Models Get Lost in Attention cites this paper.

Large Vision-Language Models Get Lost in Attention See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:26:10.066217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-08T11:54:01.224588Z digest=sha256:28c35acd07b0a8d5ea40182a682e1e928c7dbac5880e6c50ac43f07ba5a97608

Observation a3d5c165-df40-40d4-9ac9-b47f11a0ca60 · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:13:13.468869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T11:10:58.225078Z digest=sha256:175a3f4291e8d72615afa93a1d970d1f459218e4b42d2305880329c674c90394

Observation 8f45ea5d-edad-4a0f-9064-b4ea98aaa6fe · inbound

RAVE: Re-Allocating Visual Attention in Large Multimodal Models cites this paper.

RAVE: Re-Allocating Visual Attention in Large Multimodal Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:55:48.355460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T18:36:56.248838Z digest=sha256:54031f20fc66e59897033b97d1d5db8fb4a08017823742db2e4dd0c365664a8f

Observation 145b7c44-0cda-4d95-90c0-ff50baaa55d8 · inbound

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues cites this paper.

MLLMs Know When Before Speaking: Revealing and Recovering Temporal Grounding via Attention Cues See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:14:42.456705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T07:13:43.716510Z digest=sha256:ae5cea391ef81802d7a9d65753dca8cd7188c5ef5f3f82711941f6b7f76635a5

Observation 941f5569-b466-421a-8c5b-ffb153b9be89 · inbound

Inference Time Optimization with Confidence Dynamics cites this paper.

Inference Time Optimization with Confidence Dynamics See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:14:37.619616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T11:12:46.757296Z digest=sha256:8499ec5c08d9b9b2baf8a842078e09c2832bada0034a828b3c05948c8a4d38e7

Observation f7cf9279-2f32-40c9-be26-fe63dbc86178 · inbound

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning cites this paper.

Addressing Exacerbated Attention Sink for Source-Free Cross-Domain Few-Shot Learning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:34:01.800634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T22:28:35.155099Z digest=sha256:9e25c2a5da439e75c11c22492566a11fb8022487136a9c64e16622eabf82350b

Observation a41cd955-68d3-41d9-a696-11c4074591a9 · inbound

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models cites this paper.

When Graph Tokens Sink: A Mechanistic Analysis of Graph Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:56:27.409378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:27:09.779776Z digest=sha256:58d1755fff8a4ba9ee2f40eed12c477dba2f083ede89395488c79f6d7a4d6855

Observation b37d17ce-7aef-4aaa-8d64-58ba4f066f4e · inbound

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads cites this paper.

Mechanistic Insights into Functional Sparsity in Multimodal LLMs via CoRe Heads See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:46:56.644022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T01:54:13.927736Z digest=sha256:ff3ab68467be356121c1a5a37b02c538bfde82d21a7563dc3e3e16b9e2c3a51d

Observation b039fe84-d7ec-463f-a462-d74ffb2c4031 · inbound

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning cites this paper.

Reason Twice: Segmentation via Candidate Discovery and Comparative Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T00:47:30.634436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T17:01:13.745646Z digest=sha256:238b84101d8d03f511660ab6089ed64ef9f086c19054f5eb025acbf1e1906af0

Observation afb174be-9df7-440c-af56-d881b879226d · inbound

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs cites this paper.

From Senses to Decisions: The Information Flow of Auditory and Visual Perception in Multimodal LLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:57:32.316937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T16:12:53.387567Z digest=sha256:7d49fdde2007aafd627a897b6003dab5c23a56c7e8df60f25418217a373e2e73

Observation f663055d-df4d-4e1e-b486-d9cf7ee7eb4a · inbound

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models cites this paper.

Reroute, Don't Remove: Recoverable Visual Token Routing for Vision-Language Models See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.108386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T09:35:24.118536Z digest=sha256:d60452a03ae4a3ff005b0d98c8edcd66e1ab03b3e832719b2a1c17e111136b80

Observation d90ea871-600b-40cf-900d-1c1f81e05318 · inbound

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression cites this paper.

Last But Not Least: Boundary Attention CalibratiON for Multimodal KV Cache Compression See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:17:49.042253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T10:22:29.131570Z digest=sha256:f5694acaa6c6b4beab7f0f1a5455445f7b51c5edf1015ad960671e8e617358de

Observation dbfcec07-4c4c-43c2-92b9-4538d13d9812 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:19:31.048584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:97b2e55d8d476f3b87b9694e510b6938263a539e4199bbf3ee969d77b6ebcbd3

Observation 5082a377-8813-4fe3-ac05-7b19dd3bb3f9 · inbound

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context cites this paper.

VisReflect: Latent Visual Reflection for Fine-Grained Perception in Long Visual Context See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:04:21.286659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T06:01:49.904752Z digest=sha256:020186237b2ff952bcaa0837562aca461c0202c6cccc9a41538dc0f4e946ecce

Observation 9d60cd50-a443-4612-955b-80f85d85c4cc · inbound

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs cites this paper.

ADAPT: Attention Dynamics Alignment with Preference Tuning for Faithful MLLMs See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.486668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T06:51:01.371390Z digest=sha256:e4fd18b78ee6547f6d7005c82d7941832d0a0c13e25a32b54ca8b694817adb03

Observation 669815dd-0801-42ec-8b36-de554e29f9db · inbound

Information-Regularized Attention for Visual-Centric Reasoning cites this paper.

Information-Regularized Attention for Visual-Centric Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-02T15:17:07.247821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-02T15:12:43.475802Z digest=sha256:746dfd78c3b6609c1eb67ccf94aadbab3384497f0e50d3c5a9bb80c09185db9c

Observation 6ff1c143-7362-4272-bccc-416129289b67 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T05:37:19.632869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T05:37:19.632869Z digest=sha256:49b2975c3ff9e81ab7632dccdbb452acfb86eb068888f0c095eca3b37d4e727c

Observation e2cba484-7271-45db-ada6-1c7465442908 · inbound

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning cites this paper.

The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T07:01:19.935203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T07:01:19.935203Z digest=sha256:9ac7cdfb9e0dab9d2acb33edaeccbdb627168ead53530255a00a2cd7ca9f40c4

Observation dfaf8944-d82a-4aba-bcd1-997dde3b7c7d · inbound

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings cites this paper.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.730157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.730157Z digest=sha256:84383edfbd87fcfcbbc3dd5e79e5e49c64f12c1050223711a9b352c9e46842ce

Observation a9d423ba-0d22-4dc5-b4a2-33fed6b76e9e · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:35.906973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:35.906973Z digest=sha256:df53097303f031cc27691d843ddad8e3ed0db85c117e9b9f68c5d5afd19ee573

Observation b5b66a19-8b7a-43b8-b13f-0909d0cc1a4a · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:54.711000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:54.711000Z digest=sha256:48651fa1ba32e83bdaf94646f3fe64c0c89c0dfa5aa2e6ae45d32565bf49e0ef

Observation cd3f7e0f-08b8-4be6-bc56-4994d17e9b8b · inbound

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles cites this paper.

Not All Redundant Tokens Are Alike: Analyzing Visual Token Pruning through Token Roles See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:28:37.236788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:28:37.236788Z digest=sha256:2227c8605cbcf826d1bee620197523f99af4377369970b29096e8a535e742b2a

Observation bc3af201-0a87-4966-a990-2f6c2f8c9e5d · inbound

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination cites this paper.

Same Attention, Different Truths: Put Logit-Lens over Visual Attention to Detect and Mitigate LVLM Object Hallucination See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T10:47:43.957166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T10:47:43.957166Z digest=sha256:14add4442c63d8bf23575fabc7480a3d5d2e1f31b874509c1c649c8e2c018c6f