Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:29:25.650175Z
Paper Citation Record · LEDGER
As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2604.12896.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-10T15:29:25.650175Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
39 of 39 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4d716c47-5f46-4c19-8567-f92b4aecbce5 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Per- ception tokens enhance visual reasoning in multimodal lan- guage models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 303a81e6-0a18-414c-a1da-9dcd0f3f3662 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation af90f8ee-3e68-4ed5-813e-664d3890c303 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bdb2de89-3591-442e-b628-56222da10caf · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs MMFactory: A Universal Solution Search Engine for Vision-Language Tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e2acb4af-17dc-4eab-ab7c-239ce84866ee · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs GRIT: Teaching MLLMs to Think with Images
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0c876103-aa47-44f8-b5d4-3bc3c96474ea · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Hidden in plain sight: Vlms overlook their visual repre- sentations
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 23821e26-51d2-4539-a125-a4e668480c5e · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 61f3ea5c-202e-48f5-ae94-beaef5b02923 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Blink: Multimodal large language mod- els can see but not perceive
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 248ff6d7-2c0b-40f5-9547-a92a613a5b9b · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Visual program- ming: Compositional visual reasoning without training
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df35a4e1-385a-4265-af30-5d1c0a6d1b66 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vi- sual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Informa- tion Processing Systems, 37:139348–139379
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation bf45f232-5d31-45fd-abcb-9c0ea52bc4ee · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Zebra-cot: A dataset for interleaved vision language reasoning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation fb98abb0-49e8-4d7d-a333-73c31baa0d74 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Llava-plus: Learning to use tools for creating multimodal agents
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation df645dfa-8df4-4b18-811b-69b3acbd51a2 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Grounding dino: Marrying dino with grounded pre-training for open-set object detection
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 767e7b78-2c8d-4e9e-a64c-ca61b4e04aa5 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Latte: Learning to think with vision specialists
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation dcca2f04-7181-4814-8539-f2377c8e7737 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Gpt-5 system card
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 60abb331-28db-486e-a432-73f4ec20ddb8 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ef04fe77-52e5-47ce-b185-40b53d1cc8d0 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Grounded Reinforcement Learning for Visual Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad68073c-d313-45b0-b4ff-760dc9255775 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Loftr: Detector-free local feature matching with transformers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 1400c1a6-5de6-47c5-8590-c7a4e54a5dae · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vipergpt: Vi- sual inference via python execution for reasoning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9519956b-d8d7-4f96-8b4f-8f5bcae686fa · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Emergent correspondence from image diffusion.Advances in Neural Information Processing Systems, 36:1363–1389
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 08a21b56-9925-4bf3-a657-0c4a718ce2b4 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Tulip: Contrastive image-text learning with richer vision understanding
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 75506738-098f-4cc5-9006-a381ef6e2830 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Qwen3 technical report
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 63baa1e1-b5ec-4784-bcf4-0d1ab3d8bb23 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Raft: Recurrent all-pairs field transforms for optical flow
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 0cac81a4-60f5-43ca-ada6-a04a7df1ce1a · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 5a86a0d0-18d2-4bec-9c6e-3b84c97c2383 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 47cbdb7e-ffd7-4a73-95e5-b9f1ac60d283 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation edd76505-38a9-4b5d-8cb5-af12a9118c6f · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation cd8cd4ee-0ba8-46f9-be89-3e322b5230d6 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Depth anything: Unleashing the power of large-scale unlabeled data
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4c6ad6c9-fdf3-43df-8778-46a1964bd261 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation ad7551bf-0755-407d-bbea-79925e110763 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Introducing Visual Perception Token into Multimodal Large Language Model
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 4a740743-3438-4ae6-9ad0-ab9a1b75d28e · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 65ead13e-0d5c-4b26-bcff-61e6f648efcc · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation e796cd22-198b-4d2f-bedd-16dcb0c9a8a4 · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Thyme: Think Beyond Images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 9d74e71e-3af3-426f-a6cd-fa668f1aaefb · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vipact: Visual-perception enhancement via specialized vlm agent collaboration and tool-use
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 7c0852ea-11b5-4a98-be9c-18e7298cf73c · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Reinforced Visual Perception with Tools
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 827ec209-92dc-4c38-ae08-0b7fdcfc99bd · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Mainly, we provide samples of prompts for both frontier and open-source MLLMs
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d0ea9244-4b20-4433-b8aa-5651e641f6c2 · outbound
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation d94cf8ab-521a-4cb9-8777-7e17d42d456e · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs 5.1 we discussed the quality of visual interpretation of current MLLMs
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
Observation 89665866-e434-43de-820b-3ef85228df5b · outbound
Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Unresolved cited work
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.
No inbound Pith citation observations are available.