Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:03.098852Z
Paper Citation Record · LEDGER
As of 14 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 31 inbound Pith citation observations for arXiv:2501.05452.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T21:17:03.098852Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:12:03.512804Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T19:20:07.268611Z
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 551b367e-c72b-4d5c-9396-fea8378b5a6a · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890e7e32-0475-4032-b9ab-f5dd97917b89 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Vip- llava: Making large multimodal models understand arbitrary visual prompts
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ae235e4b-6838-4c9c-b6bc-a9ca53e6c85f · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Bigtable: A distributed storage system for structured data
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1bc0f498-95fd-4cfc-9191-e03ed9842b63 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TabFact: A Large-scale Dataset for Table-based Fact Verification
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f35a8280-5f57-40c7-be88-2dc5ef7ac549 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding HybridQA: A Dataset of Multi-Hop Question Answering over Tabular and Textual Data
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 129445eb-5a02-45ef-b9d8-25ca240451ea · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual chain- of-thought prompting for knowledge-based visual reasoning
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 752f8bec-c5f1-4de5-ac7e-268977014af5 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Selective attention and the organization of visual information
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fd405925-dcba-4488-a58c-c3c3c3f77105 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48de0c73-444f-4ae3-9837-37c80ec70877 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding The cambridge structural database
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 12fe13e4-b780-485f-bd09-0cacd42b10c9 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual program- ming: Compositional visual reasoning without training
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation aba33c52-2cba-43ff-8195-ce4c1902d980 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding ChartLlama: A Multimodal LLM for Chart Understanding and Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c03a0a0e-a2ae-40c3-b22c-1bdda52436ff · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding LoRA: Low-rank adaptation of large language models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 6accc76a-4a80-4525-a1cb-aae5763a6e31 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual Sketchpad: Sketching as a Visual Chain of Thought for Multimodal Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e4d4111-8e2e-4203-b519-db631c3aa707 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Selective attention
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 5a1c74f1-a1ed-437f-8949-5ad9bc45fd2b · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Dvqa: Understanding data visualizations via question answering
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 10d939ae-e693-44ef-a778-5f7e649cadc1 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe7bad8c-76f4-4d13-b236-24af8e57085e · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Semantic-SAM: Segment and Recognize Anything at Any Granularity
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6835a0-9c66-440e-b350-7582cd28631e · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart Derendering
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71447ef2-9335-42e7-9fa4-9212da114fbc · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Deplot: One- shot visual language reasoning by plot-to-table translation
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4a436c11-afc5-4582-88d6-53512bb7049b · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Llava-next: Improved reason- ing, ocr, and world knowledge, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 22f212fe-bd87-48cc-bd6a-979246823546 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding LLaVA-Plus: Learning to Use Tools for Creating Multimodal Agents
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36634454-de3a-4607-9329-e070e89430c8 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d041773-43db-41e4-a747-21eae5fc0896 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Ocrbench: On the hidden mystery of ocr in large multimodal models, 2024
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b31274e7-bc60-47e9-9fb3-865ffedb965a · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5664581-51f7-4884-9d9c-5a0593773aab · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding UniChart: A Universal Vision-language Pretrained Model for Chart Comprehension and Reasoning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 434f9a2e-c558-4962-8fc9-049fdb0a98f2 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Gpt-4 technical report, 2023
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eaab08c5-a119-467b-be8f-8388fc6fce3a · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Compositional Semantic Parsing on Semi-Structured Tables
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228911d1-b60e-4c00-8ea7-d7b70e3c56c7 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Space and selective attention
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0b96053f-5827-4957-a6b3-d76a212ded52 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Visual cot: Advancing multi-modal language models with a comprehen- sive dataset and benchmark for chain-of-thought reasoning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 2b3046a8-7b7c-4207-a61f-eaf32d4fbaf7 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding What does clip know about a red circle? vi- sual prompt engineering for vlms
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 96b638de-b646-417e-a14a-a74937bfeef9 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Vipergpt: Visual inference via python execution for reasoning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c1588615-5afe-4728-8fb3-8b12e9b238c5 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Gemini: A Family of Highly Capable Multimodal Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db01fd5-aa86-43e9-b497-7b07c4661d31 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7019453c-5e3c-4b32-8505-c6115d5c482c · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Chain-of- thought prompting elicits reasoning in large language models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation fc8289f2-9294-4a24-b9b1-bbdda74fc2e5 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ca0e2e-60f4-44c2-9e1a-ab94a3e89909 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aad3ddb-f5d3-4ee6-bb04-eeae1003bac4 · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Depth anything: Unleashing the power of large-scale unlabeled data
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8e309efd-9dba-4b97-92e4-195337ba7a2c · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding Multimodal Chain-of-Thought Reasoning in Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a46cecd2-4a68-4e7f-983e-da1af320d80a · outbound
ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding TAT-QA: A Question Answering Benchmark on a Hybrid of Tabular and Textual Content in Finance
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e425596d-f5a0-4bab-adea-2f6a8358a6e1 · inbound
Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 123
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 77c84b54-142c-4262-9e70-9a820e655541 · inbound
Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement Finetuning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ea2b8a-629e-4183-9890-b255564e406e · inbound
Grounded Reinforcement Learning for Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 22bcc49e-abe1-4239-a509-8b6c2363d5ac · inbound
ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc20143-f1e1-43d1-bd7f-deb87faafb28 · inbound
TableMoE: Neuro-Symbolic Routing for Structured Expert Reasoning in Multimodal Table Understanding ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be5a5374-0176-4028-a3ab-e3510564f0d7 · inbound
Learning Only with Images: Visual Reinforcement Learning with Reasoning, Rendering, and Visual Feedback ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbc72a80-308f-4e3c-b271-e62f4bef7e53 · inbound
Empowering Nanoscale Connectivity through Molecular Communication: A Case Study of Virus Infection ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad25cc20-36be-4892-b653-c9ffa711c9c5 · inbound
Empowering Multimodal LLMs with External Tools: A Comprehensive Survey ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 130
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e9a0b5b-0735-4ee7-9491-a579f047b008 · inbound
Visual-TableQA: Open-Domain Benchmark for Reasoning over Table Images ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 998508ec-f91f-4376-ac1b-79235175a383 · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f37e4270-47d3-4738-ae3b-4ba2172e483d · inbound
DeFacto: Counterfactual Thinking with Images for Enforcing Evidence-Grounded and Faithful Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8a1e917d-b360-4c8a-abcc-d04eaf46d0a8 · inbound
Latent Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 219fe1a3-7932-4d13-b87d-31de3c9157f6 · inbound
Training Multi-Image Vision Agents via End2End Reinforcement Learning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 03055f22-2501-4018-be04-b3f9247b503c · inbound
Reasoning Within the Mind: Dynamic Multimodal Interleaving in Latent Space ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation b0158267-52e2-4167-bac1-f521860f18d4 · inbound
Fully Spiking Neural Networks with Target Awareness for Energy-Efficient UAV Tracking ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3904687a-9598-41e6-9b8f-70d966161ec2 · inbound
Learning to Focus and Precise Cropping: A Reinforcement Learning Framework with Information Gaps and Grounding Loss for MLLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 1ee1dbbd-ac54-421e-b3d6-e99745b3c56d · inbound
TableVision: A Large-Scale Benchmark for Spatially Grounded Reasoning over Complex Hierarchical Tables ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 117371bc-87f1-486e-a6d4-b5027cdfda2b · inbound
Decompose, Look, and Reason: Reinforced Latent Reasoning for VLMs ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 0ef29106-7f0d-4c59-9e7a-af3a3bb4a19c · inbound
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 4c4725e2-eccb-4442-92bb-9d091f470937 · inbound
MedLVR: Latent Visual Reasoning for Reliable Medical Visual Question Answering ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f704a458-45c9-43e1-8da5-c1b1fe9776f4 · inbound
Meta-CoT: Enhancing Granularity and Generalization in Image Editing ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation dff89b1a-891e-48ca-b7e3-46cb9a7a0c52 · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation f1b918f8-0771-4b8f-aeca-509014fe7a7c · inbound
Vision-OPD: Learning to See Fine Details for Multimodal LLMs via On-Policy Self-Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation c0f0b526-4d84-449c-935b-dc8265e2006d · inbound
STORM: Internalized Modeling for Spatial-Temporal Reasoning in Video-Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation d66e4bf6-435c-4712-bc60-d69a1260106f · inbound
DeepLatent: Think with Images via Parallel Latent Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 34188065-5041-4a10-9a42-0215a8c9592f · inbound
Dive into the Scene: Breaking the Perceptual Bottleneck in Vision-Language Decision Making via Focus Plan Generation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 39696f71-5678-4cc6-9ead-75adaa80cd11 · inbound
Latent Visual States for Efficient Multimodal Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation ad5ce396-1ef0-4591-b632-ed4e207cee02 · inbound
V-Zero: Answer-Label-Free On-Policy Distillation with Contrastive Evidence Gating for Fine-Grained Visual Reasoning ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation caddb4f2-0306-4499-9eee-e6e140c10e32 · inbound
From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 152
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.
Observation 8fa4e668-6590-46fe-9b7c-0c84de0cd499 · inbound
FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d35572-6fcf-4aa5-89c6-796679fd1500 · inbound
VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation ReFocus: Visual Editing as a Chain of Thought for Structured Image Understanding
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.