Pith. sign in

Paper Citation Record · LEDGER

Attending to Multimodal Generation One Token at a Time

As of 6 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2607.03738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03738 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T00:17:46.528896Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 08fa7a93-aaae-46e4-8fef-3656a45d7d9d · outbound

This paper cites Possible Principles Underlying the Transformation of Sensory Messages.

Attending to Multimodal Generation One Token at a Time Possible Principles Underlying the Transformation of Sensory Messages

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:fb069a249667058c5f2391075f0c8981be287e7e475643967a88d37fa5b92326

Observation cbaa9ab6-6293-4a71-95f3-0b4c700d7d3f · outbound

This paper cites DEX-AR: A Dy- namic Explainability Method for Autoregressive Vision-Language Models.arXiv preprint arXiv:2603.06302, 2026.

Attending to Multimodal Generation One Token at a Time DEX-AR: A Dy- namic Explainability Method for Autoregressive Vision-Language Models.arXiv preprint arXiv:2603.06302, 2026

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:4c6bcd9e28d0d25cfac01108dbbac4f1011c9fb8d0f49f060a3e1d392dad10c0

Observation e356a26e-3e6a-4388-bbe1-7c7bb08cb03a · outbound

This paper cites Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation.arXiv preprint arXiv:2509.22496, 2025.

Attending to Multimodal Generation One Token at a Time Where MLLMs Attend and What They Rely On: Explaining Autoregressive Token Generation.arXiv preprint arXiv:2509.22496, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:25e7f9f5957525394cd3c34d9388e38c5172de8578319849921a34e5ebe9fea9

Observation 328f21a5-5036-48f6-8417-501111917bd9 · outbound

This paper cites What Does BERT Look At? An Analysis of BERT’s Attention.

Attending to Multimodal Generation One Token at a Time What Does BERT Look At? An Analysis of BERT’s Attention

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:d33a0b766a0a3121aa43186c03e3086f44470ce1eb881ba66bef3ef8576e24b8

Observation 01ef25fe-f403-4a05-b5cd-fcbf9b1d5c5d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Attending to Multimodal Generation One Token at a Time Training Verifiers to Solve Math Word Problems

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:003d850f9665d4d0debd978eb99d77d1aa9892f63c6acae0788d4deda6b9ee1b

Observation 4e76bd23-0bf3-48ce-85b6-1d6c4a37c165 · outbound

This paper cites Flashattention: Fast and Memory-efficient Exact Attention with IO-Awareness.

Attending to Multimodal Generation One Token at a Time Flashattention: Fast and Memory-efficient Exact Attention with IO-Awareness

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:44cebb987db8fe3b546adea1d77aab494876c9a18aafa98e3022802d5aa9737b

Observation 1cc264bf-6317-480b-b40a-ffccdf5af6a3 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition At Scale.

Attending to Multimodal Generation One Token at a Time An Image is Worth 16x16 Words: Transformers for Image Recognition At Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:873a3e7ecf31eacad3328a7fc5dc61c76d9153f96309d966d2d21897b4f684b9

Observation e01c8e83-32e3-4ec4-8252-5e6c76fe734e · outbound

This paper cites A Mathematical Framework for Transformer Circuits.Transformer Circuits Thread, 2021.

Attending to Multimodal Generation One Token at a Time A Mathematical Framework for Transformer Circuits.Transformer Circuits Thread, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:e799b2dd23c1821b6f595ad7a2509b6f719e6eb94e3b823cd5de368031529800

Observation a51830b2-a3c7-4472-99f1-d8d35a38e971 · outbound

This paper cites Toy Models of Superposition.

Attending to Multimodal Generation One Token at a Time Toy Models of Superposition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:b79709a48f0d80c3d767fcf0e3b71828b71c4f063b93ede596b0ed1c7a14301c

Observation 255440bb-5c44-433f-9068-34e0d571fefb · outbound

This paper cites Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration.

Attending to Multimodal Generation One Token at a Time Mitigating Hallucination in Large Vision-Language Models via Adaptive Attention Calibration

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:db2f3160a450cbf782bec5b919c722183455afaf2c1ea3f87001fbed6ab08b88

Observation 40dc75b4-0d16-487a-91fe-7c6b987a25d7 · outbound

This paper cites Dissecting Recall of Factual Associations in Auto-regressive Language Models.

Attending to Multimodal Generation One Token at a Time Dissecting Recall of Factual Associations in Auto-regressive Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:74dfaa6c05763ea04275ea40687b09f4c155ef2f94e39c7079b0898379ce689e

Observation 91ac1e6f-8735-46c2-916a-6a7bb33a924a · outbound

This paper cites LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts.

Attending to Multimodal Generation One Token at a Time LLMSteer: Improving Long-Context LLM Inference by Steering Attention on Reused Contexts

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:c39176599098313409692af6ed045f77605c0fa3facc91ae12ed020b60d13fc9

Observation 218f5833-c688-4f6d-b7e9-af96fcff062a · outbound

This paper cites How Do Vision-Language Models Process Conflicting Information Across Modalities?.

Attending to Multimodal Generation One Token at a Time How Do Vision-Language Models Process Conflicting Information Across Modalities?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:eb2ac82040686d69318f119125349a86f270b647aa45229d8f78f32adea2d589

Observation cf0ca37f-581a-487d-84b3-d3a5906274c6 · outbound

This paper cites Opera: Alleviating Hallucination in Multi-modal Large Language Models via Over-Trust Penalty and Retrospection-allocation.

Attending to Multimodal Generation One Token at a Time Opera: Alleviating Hallucination in Multi-modal Large Language Models via Over-Trust Penalty and Retrospection-allocation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:3cb7521e5f172bc9e72b26fda6b03179a624e363f8c33810d781efe7a58b905c

Observation 8ef6b9de-9153-46ef-85ef-45280e840622 · outbound

This paper cites The Platonic Representation Hypothesis.

Attending to Multimodal Generation One Token at a Time The Platonic Representation Hypothesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:82374d0389413872cc61f7d0d146fb4bb7ddb43bb45b4698f5f4b3a2640c854c

Observation 96b94612-2c44-4e77-9062-bd2694f9f545 · outbound

This paper cites What’s in The Image? A Deep-Dive Into the Vision of Vision Language Models.

Attending to Multimodal Generation One Token at a Time What’s in The Image? A Deep-Dive Into the Vision of Vision Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:d41540197740a2d9ee0de9632fbb4e36c3159af391d4a04afb8946c9408a287c

Observation b215ee15-3dfa-4098-8ee6-3d945a17fa6b · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Attending to Multimodal Generation One Token at a Time See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:7a130d00c1658049ea0782fa13497ce8bc3cb783250051f4d56d55cfe5bc0c45

Observation 7fffb951-82e2-4faf-b8b3-305bea08cfc9 · outbound

This paper cites The Open Images Dataset V4: Unified Image Classification, Object detection, and Visual relationship detection at Scale.

Attending to Multimodal Generation One Token at a Time The Open Images Dataset V4: Unified Image Classification, Object detection, and Visual relationship detection at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:46c76d6e3eac36e3780f21ad921c4757fa899cd01854c9c65d8969b8383a8fb3

Observation d84c590e-5e3c-4800-bf06-386a799c9d56 · outbound

This paper cites LLaV A-OneVision: Easy Visual Task Transfer.

Attending to Multimodal Generation One Token at a Time LLaV A-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:b616df6cae071dcfb0bacbfed2b4791ecf9282005de7e715904edbd8572430f3

Observation 977d2468-8378-4443-8248-fc8081ad1199 · outbound

This paper cites Dynamic Token Reduction during Generation for Vision Language Models.

Attending to Multimodal Generation One Token at a Time Dynamic Token Reduction during Generation for Vision Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:979ab4fa016f0ce886b51614535634c2a181e76698ff8ee34122e941ae6a115b

Observation 53c7fdd9-37a2-492c-a278-24b313ed3881 · outbound

This paper cites Paying More Attention to Image: A Training-free Method for Alleviating Hallucination in LVLMs.

Attending to Multimodal Generation One Token at a Time Paying More Attention to Image: A Training-free Method for Alleviating Hallucination in LVLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:a3a2beaf381224eab862ebbe83bc2d3eba01497bafbba47d52cb786cbbde01d7

Observation 6cc6c3e1-9ce9-42d9-88c1-e04dfdfd7a77 · outbound

This paper cites Vision-language Models Create Cross-modal Task Representations.

Attending to Multimodal Generation One Token at a Time Vision-language Models Create Cross-modal Task Representations

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:bf3c2d2b522df9136b88326a1e159556e3117178e3eb66ff41966a31c62163ac

Observation d0593d9f-a0d8-4701-90f4-ee1c85f081b7 · outbound

This paper cites Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs.arXiv preprint arXiv:2602.15556, 2026.

Attending to Multimodal Generation One Token at a Time Revealing and Enhancing Core Visual Regions: Harnessing Internal Attention Dynamics for Hallucination Mitigation in LVLMs.arXiv preprint arXiv:2602.15556, 2026

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:9404308c0588b408dd23fac09b4908670d5715175e5cf5af201b172aece2cf8d

Observation 9dd6ebce-b692-414a-9650-585960a38667 · outbound

This paper cites Chartqa: A Benchmark for Question Answering About Charts with Visual and Logical Reasoning.

Attending to Multimodal Generation One Token at a Time Chartqa: A Benchmark for Question Answering About Charts with Visual and Logical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:7494dd8aea7bbe0f1eeba3a997435a1b8bde822985642633d6791ab9ebc3a669

Observation 8ee55ee0-5c34-45fb-b29c-d0f3e4c9a057 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Attending to Multimodal Generation One Token at a Time Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:d1e3900d124a8060b4b51fc2d535321c574b9ab5d1b4db836661022b75f85632

Observation 064b8b0c-de79-4c92-903a-a1c03b51a728 · outbound

This paper cites Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs.

Attending to Multimodal Generation One Token at a Time Same Task, Different Circuits: Disentangling Modality-Specific Mechanisms in VLMs

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:cdb5cb90d8af95393fc373e7695ffb2babbbeb894574880f74129d3709744d30

Observation 9e7748e5-bde2-4454-b870-f7bdc1c5caae · outbound

This paper cites Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions.

Attending to Multimodal Generation One Token at a Time Token-wise Decomposition of Autoregressive Language Model Hidden States for Analyzing Model Predictions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:50c141387952feca3a902047265ce950de77901d091b33623b96735494f924f3

Observation 6db87029-9da1-41f3-825e-452ed1f597ec · outbound

This paper cites Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-language Conflict.

Attending to Multimodal Generation One Token at a Time Mixed Signals: Decoding VLMs’ Reasoning and Underlying Bias in Vision-language Conflict

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:160a683f30600c31360dc68d2474ea52031104cb890ef309dd1ff5d9067cc6bb

Observation 70e14a40-542f-4983-8951-cb37a91a1d47 · outbound

This paper cites Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding.

Attending to Multimodal Generation One Token at a Time Seeing Far and Clearly: Mitigating Hallucinations in MLLMs with Attention Causal Decoding

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:d3a4c51d86748339b1f15bb29f01f7da0a2de75580da085d61534b7684f9b439

Observation 9b5e3a1b-0c97-4b16-8560-b3fd88e668f1 · outbound

This paper cites Attention Is All You Need.

Attending to Multimodal Generation One Token at a Time Attention Is All You Need

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:e5ddb158f87de3583cbe7b2bf5a5f6f540d424a8177bf869af4473e790df8932

Observation dbe5126e-3d39-4a7b-8488-04d39a2ab45e · outbound

This paper cites MLLM Can See? Dynamic Correction Decoding for Hallucination Mitigation.

Attending to Multimodal Generation One Token at a Time MLLM Can See? Dynamic Correction Decoding for Hallucination Mitigation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:ecfc7ac65185bd52f7a73674b72bddeb20b1a5a9cfc30b5b4c2a70bebe9f33f8

Observation c0ffc56c-43c1-4b06-8cac-18bcca0798cb · outbound

This paper cites ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM.

Attending to Multimodal Generation One Token at a Time ASCD: Attention-Steerable Contrastive Decoding for Reducing Hallucination in MLLM

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:8af5b4bf6e06a011a87b129df3fff2315a527e5fe77e0cbd51de10048da52845

Observation 0493d441-b936-4d9a-b31a-c812ae6db9ed · outbound

This paper cites Measuring Cross-modal Interactions in Multimodal Models.

Attending to Multimodal Generation One Token at a Time Measuring Cross-modal Interactions in Multimodal Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:99a2693c34ac1e91b0df6668746bac6ff352934d410a25d13e5c5ec4dfdd279b

Observation 4b0ac126-b5b9-4f21-9d6c-567fc5234c06 · outbound

This paper cites Qwen2 Technical Report.

Attending to Multimodal Generation One Token at a Time Qwen2 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:813e0a2e5c736dc47adb3ce684d5c53c294944cd597deffb47f030e7dfed7d92

Observation 5da9e978-3edb-434d-b4e7-85b94844be78 · outbound

This paper cites Qwen2.5-VL Technical Report.

Attending to Multimodal Generation One Token at a Time Qwen2.5-VL Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:35d3f7f5f15b3d76ee135b07d19fb41275ba6a3f3daf0db8dcb58fba54135b38

Observation f0c1f820-2f83-4de0-b8a8-8d8680dd46a4 · outbound

This paper cites Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference.

Attending to Multimodal Generation One Token at a Time Lifting the Veil on Visual Information Flow in MLLMs: Unlocking Pathways to Faster Inference

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:12955d72b21806b80e6c4a3f5861ef9a44dbbb087afce9ea815f90721a912985

Observation f3b96e57-7054-4b3d-900c-b5529aa7bfbf · outbound

This paper cites SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification.

Attending to Multimodal Generation One Token at a Time SAVAA: Mitigating Hallucinations in LVLMs via Step-wise Adaptive Visual Attention Amplification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:9f32a0b29c6db1f79ce83b93af8733ea00fa8c9de183489a143fad72a44aec73

Observation 714997e9-62f3-4514-ac7d-ab77b2950a31 · outbound

This paper cites Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs.

Attending to Multimodal Generation One Token at a Time Tell Your Model Where to Attend: Post-hoc Attention Steering for LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:0b895c4ac5138280f90bee0fc479491b889cdb00fc79f4d06299e0cb258530c4

Observation 392a7a09-f0ce-4e1b-897a-0b485d5922c2 · outbound

This paper cites Adaptinfer: Adaptive Token Pruning for Vision-language Model Inference with Dynamical Text Guidance.

Attending to Multimodal Generation One Token at a Time Adaptinfer: Adaptive Token Pruning for Vision-language Model Inference with Dynamical Text Guidance

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:a88d240ffb1fc135ab3d41d3a599292e5f84ea076565cf81581c102140c9dbb6

Observation 04d4c8a1-6d05-4159-82fe-6e40f215564c · outbound

This paper cites Cross-modal Information Flow in Multimodal Large Language Models.

Attending to Multimodal Generation One Token at a Time Cross-modal Information Flow in Multimodal Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:438f8bd7031f8754e3eb0bcd20304c9cff90522aa269ebdff94ffc646576beec

Observation 4e904484-1923-4b72-919f-8950b8b4a664 · outbound

This paper cites The fruit is.

Attending to Multimodal Generation One Token at a Time The fruit is

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:247f9d35d36d0fb73a04a8dbe0fcc9868168ef5a2c4886c063d9f6e29e68edca

Observation e741380a-59f3-4af8-a066-82aaf0231b8d · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:8bff4fd5e06b328ddd4a9d7b77eb49531ef4442ac8b9402f67a6d23046b297ea

Observation 6f1111c9-50c4-410e-b569-81361cbbf630 · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:31a5798f975978b10ea737207d1eafccb3ea6166db4c56e31169d8aa72a54de3

Observation b4340adb-fe09-4a85-a1da-bfbc72f0dbc7 · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:774df6f39a0e54b72a1d56077712cab638fd540fa50df46667f583b930d897da

Observation befe447d-60b4-4bbd-938a-f0887912f8de · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:c10b5a9751ac1ee839416b48eb03e6548c4575e10374685b8fa5ea523696f961

Observation e018c8f3-4fc9-485a-9588-9aa61c6ee5ab · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:80455d9b78689b3586d99623f35ea7ab85b1717557cc4150bdab561e967fe325

Observation d61ba245-aea7-4cd5-ab76-307396136305 · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:ddebbc79d28b4db1237892612b910d8344d4039b6cf4ee9e5dc3017328efd83e

Observation 94695060-afcf-4b9d-9ff2-36289ce43d49 · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:58b755135e2c37bf5490ca5ca895a19715fc4239bde4390471c5a612e909ab1b

Observation 5d3173e9-799c-45d7-a22f-fb2a4f7fb759 · outbound

This paper cites an unresolved cited work.

Attending to Multimodal Generation One Token at a Time Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:e0fd0101839c6b1fe048eec868334080fe3731ac3277eaf020da2174c7c0ba8f

Observation 30d3a132-9d79-4ac7-9673-2ada6ed01cfc · outbound

This paper cites I see no fruit... it is a cherry.

Attending to Multimodal Generation One Token at a Time I see no fruit... it is a cherry

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:b7f66e32815f0ef7f52fe8edce1257a196ad71070794e6c9ad69f1eaf8fa63ed

Observation c0bc0d37-a42d-46ca-a432-23bf3df227c5 · outbound

This paper cites fruit_answered.

Attending to Multimodal Generation One Token at a Time fruit_answered

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:ae7f38c6dfe4a5097f83bed798126713db34269f3bb863a77e71388a990c63d6

Observation fec51bd9-7d87-4682-9499-38e70ffc5699 · outbound

This paper cites True” or “False.

Attending to Multimodal Generation One Token at a Time True” or “False

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:528ae48f1cdced20079e8bd9d77cb40d1a98bd19bb1e5920c7c33e89d195415f

Observation 885cb1a8-5677-46ae-9cd5-3e8eb5b7d1ab · outbound

This paper cites True” or “False.

Attending to Multimodal Generation One Token at a Time True” or “False

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:c75a0280a1bbc94d0d21dce4db0118a8c286033ba79d3ce58b712157b6b36ed5

Observation 7db763b5-8f1e-418c-b58c-3b39625c35b8 · outbound

This paper cites visual_check.

Attending to Multimodal Generation One Token at a Time visual_check

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T00:17:46.528896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T00:17:46.528896Z digest=sha256:fc901c9031534dcf23a173ce254b4b55873a22e7266e1490e381f0e1faa453d0

Pith citing papers

No inbound Pith citation observations are available.