Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:52:36.195441Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 3 inbound Pith citation observations for arXiv:2508.20279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:52:36.195441Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-10T13:49:17.343893Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
23 of 23 outbound references displayed
External citation measurements
1
pith, observed 2026-08-05T02:28:24.338817Z
Observation 98858cfd-a862-4147-82de-a1a33a542fb1 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Wenliang Dai, Junnan Li, Dongxu Li, Anthony Tiong, Junqi Zhao, Weisheng Wang, Boyang Li, Pascale N Fung, and Steven Hoi
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 806ad925-4db3-45d3-8eaa-704092268396 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dfd39f7-2da9-4227-b077-c9105219996a · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Transcoders Find Interpretable LLM Feature Circuits
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e033c343-9192-40e3-ad5e-00400c0a5741 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Scaling and evaluating sparse autoencoders
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2656b8f-2356-46a8-8b2d-f17385bd7e63 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding What Do VLMs NOTICE? A Mechanistic Interpretability Pipeline for Gaussian-Noise-free Text-Image Corruption and Evaluation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f1d0792-cf1a-4ede-bc7a-fb26795b264d · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding How to use and interpret activation patching
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa626305-c01b-4707-b1a9-4c551602a5c9 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding MMNeuron: Discovering Neuron-Level Domain-Specific Interpretation in Multimodal Large Language Model
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 315330f3-a536-4e49-b825-6dab7b2bf359 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Quantifying Feature Space Universality Across Large Language Models via Sparse Autoencoders
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de0f6f12-f9dd-4b7b-ad90-f5d2b918a6d3 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation d7f5f1c2-47ec-4ea0-acd2-ece3ba25cc58 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21e32dec-0b15-41cc-b523-251b9567f762 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding LLaMA: Open and Efficient Foundation Language Models
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bf32ff1-16ac-4f5e-80d4-9b39ba6113a3 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18fb96d7-d4d9-42c1-b8ab-86104187e19a · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Probing Large Language Models from A Human Behavioral Perspective
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 943f7d77-626a-4680-ac2d-994b9cd7e972 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding How Interpretable are Reasoning Explanations from Prompting Large Language Models?
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 28f9fed7-a112-4e81-b7ec-e4fa5bb813b0 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding OPT: Open Pre-trained Transformer Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49d9567f-a6ef-4923-90d6-94a88b39e795 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7039e817-5b59-4788-9ce7-54c877c062e8 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6399d9f2-939c-4928-a073-41faf76a7ee8 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Qwen Technical Report
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec95d0d9-3213-41fb-bdc4-79c81d0e13ce · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97284a95-1243-46a2-99f3-ea0cc0e78305 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding A structural probe for finding syntax in word representations
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 19697ec8-1d39-4cfa-9c0f-2a9f8313185c · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding AGIEval: A Human-Centric Benchmark for Evaluating Foundation Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a797298-bb6d-4bbe-a6d1-86b7ea7bc952 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Understanding Information Storage and Transfer in Multi-modal Large Language Models
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 084f4cdc-dce6-41c9-98e9-189511aebde8 · outbound
How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding Language models are few-shot learners
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 325c9e4d-786b-4d8f-8499-f721c7ce50e8 · inbound
From Heads to Neurons: Causal Attribution and Steering in Multi-Task Vision-Language Models How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
Reference 110
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 6dbcd8c3-ad05-4cb3-9a1b-e44e89f2ed65 · inbound
The Hidden Evolution of Disguised Visual Context inside the VLM How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c976d48f-b72b-45db-a734-c8aebcdc54a5 · inbound
From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs How Multimodal LLMs Solve Image Tasks: A Lens on Visual Grounding, Task Reasoning, and Answer Decoding
Reference 198
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.