Pith. sign in

Paper Citation Record · LEDGER

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs

As of 6 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2604.12896.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.12896 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-10T15:29:25.650175Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact15
  • verified fuzzy22
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4d716c47-5f46-4c19-8567-f92b4aecbce5 · outbound

This paper cites Per- ception tokens enhance visual reasoning in multimodal lan- guage models.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Per- ception tokens enhance visual reasoning in multimodal lan- guage models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.685588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:03ee909ac16a8e7aab0889c7f855621a58d87dfb341b7f1a29f33653cde470e8

Observation 303a81e6-0a18-414c-a1da-9dcd0f3f3662 · outbound

This paper cites PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:02.000635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:97a73975ae622749ef775c6d46d5bef7b322ad458d0461a43c23f7efdb1ec049

Observation af90f8ee-3e68-4ed5-813e-664d3890c303 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:26:01.983777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:d44a2e1aa786d1afec3e2115462a6abb6be6782f4d000ef45c67dc9a3e3f2c02

Observation bdb2de89-3591-442e-b628-56222da10caf · outbound

This paper cites MMFactory: A Universal Solution Search Engine for Vision-Language Tasks.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.993879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:624a7e6c193f64bdf77d2bc43a3ddfa1a078f98d33f6c8c4b3b025fe499e3de5

Observation e2acb4af-17dc-4eab-ab7c-239ce84866ee · outbound

This paper cites GRIT: Teaching MLLMs to Think with Images.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs GRIT: Teaching MLLMs to Think with Images

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:44:46.816923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:3941df3855b591c0e10c88c8f56b0f1adaff65ce5453a77f57d35fc253311cdd

Observation 0c876103-aa47-44f8-b5d4-3bc3c96474ea · outbound

This paper cites Hidden in plain sight: Vlms overlook their visual repre- sentations.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Hidden in plain sight: Vlms overlook their visual repre- sentations

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.688708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:98f561448f44e49310419abc0071bb6ca79ffdad16afb16ff247d274b0c04d59

Observation 23821e26-51d2-4539-a125-a4e668480c5e · outbound

This paper cites Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Llmdet: Learning strong open-vocabulary object detectors under the supervision of large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.691861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:7a835e69d4f5cebc543244cc59d64a41adb840f4be3838fc86bb6f287487ec15

Observation 61f3ea5c-202e-48f5-ae94-beaef5b02923 · outbound

This paper cites Blink: Multimodal large language mod- els can see but not perceive.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Blink: Multimodal large language mod- els can see but not perceive

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.695215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:090118e94f4992f5348335ab8785879d5b6cc61f9f9272a60f3d7f9ffc69ca03

Observation 248ff6d7-2c0b-40f5-9547-a92a613a5b9b · outbound

This paper cites Visual program- ming: Compositional visual reasoning without training.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Visual program- ming: Compositional visual reasoning without training

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.698186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:ac8ee6118bec581b871525173df542675dc8760728c0ba5c394d47d25be174d1

Observation df35a4e1-385a-4265-af30-5d1c0a6d1b66 · outbound

This paper cites Vi- sual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Informa- tion Processing Systems, 37:139348–139379.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vi- sual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Informa- tion Processing Systems, 37:139348–139379

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.682501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:08fa20eb663c4f63e43b0aabfddf2218f4a11f73e3253783d65ed95deb6c556e

Observation bf45f232-5d31-45fd-abcb-9c0ea52bc4ee · outbound

This paper cites Zebra-cot: A dataset for interleaved vision language reasoning.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Zebra-cot: A dataset for interleaved vision language reasoning

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:02.008662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:8cfbfb6b4e56f08c9e260791cf095f171c5dfaed3c304b875016c6e3c2ace3b7

Observation fb98abb0-49e8-4d7d-a333-73c31baa0d74 · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Llava-plus: Learning to use tools for creating multimodal agents

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.679110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:3359c522af182a0c0408561096c26ff5b4317b9bbe3d61f31b96e6f6c4bc0239

Observation df645dfa-8df4-4b18-811b-69b3acbd51a2 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.655114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:20a7926893c0f95cdeb663b114da0f4b1c79e888ab8aeedda2d197d0cb6c69bb

Observation 767e7b78-2c8d-4e9e-a64c-ca61b4e04aa5 · outbound

This paper cites Latte: Learning to think with vision specialists.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Latte: Learning to think with vision specialists

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.670107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:f0d964e0565eb63c2fd703be3fc8f9ebc46a7a2fb45c84809e373ddfaa54874f

Observation dcca2f04-7181-4814-8539-f2377c8e7737 · outbound

This paper cites Gpt-5 system card.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Gpt-5 system card

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.629081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:6104630fce978f17ca75a35b1b2d1ab658c50507b975b65be7d4ac9c7a68d6b8

Observation 60abb331-28db-486e-a432-73f4ec20ddb8 · outbound

This paper cites Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Per- ception test: A diagnostic benchmark for multimodal video models.Advances in Neural Information Processing Systems, 36:42748–42761

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.631871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:82706f8f367e7ca2ab2cebe56c997eb9cb6817fc3b65de629469d6012a5f9a1f

Observation ef04fe77-52e5-47ce-b185-40b53d1cc8d0 · outbound

This paper cites Grounded Reinforcement Learning for Visual Reasoning.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Grounded Reinforcement Learning for Visual Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.973499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:b55019eaf720a569fdeb6c16e5869de4349dc2cd49efee7befb9c393d11f72ab

Observation ad68073c-d313-45b0-b4ff-760dc9255775 · outbound

This paper cites Loftr: Detector-free local feature matching with transformers.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Loftr: Detector-free local feature matching with transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.647855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:8bf939829598ad711ca4e6bab657274d9fa504d3952e8f0dfd95089bed916285

Observation 1400c1a6-5de6-47c5-8590-c7a4e54a5dae · outbound

This paper cites Vipergpt: Vi- sual inference via python execution for reasoning.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vipergpt: Vi- sual inference via python execution for reasoning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.658284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:225ac89e111323ccfc7e74979d7d38d2fdbc2b9f050b3e2a47945a92e8d1c656

Observation 9519956b-d8d7-4f96-8b4f-8f5bcae686fa · outbound

This paper cites Emergent correspondence from image diffusion.Advances in Neural Information Processing Systems, 36:1363–1389.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Emergent correspondence from image diffusion.Advances in Neural Information Processing Systems, 36:1363–1389

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.667044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:32e8a354c9d6bc7c497f658e4e4e23104a1c0ed6b7a9941587f3d055da660f1b

Observation 08a21b56-9925-4bf3-a657-0c4a718ce2b4 · outbound

This paper cites Tulip: Contrastive image-text learning with richer vision understanding.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Tulip: Contrastive image-text learning with richer vision understanding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.634567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:6a9a897e3ed66bed3345145a9fa64f38b61bae41079323a540efec1618356008

Observation 75506738-098f-4cc5-9006-a381ef6e2830 · outbound

This paper cites Qwen3 technical report.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Qwen3 technical report

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.637249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:b84d7141f1d61f3b25f939dea9c3d9767af9d6649d2ed43ab4c008acc8355c28

Observation 63baa1e1-b5ec-4784-bcf4-0d1ab3d8bb23 · outbound

This paper cites Raft: Recurrent all-pairs field transforms for optical flow.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Raft: Recurrent all-pairs field transforms for optical flow

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.673027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:afd0af80ede75921fb4c31b803b4e3771d73ec3c5fce2bac93e8e83873f36e02

Observation 0cac81a4-60f5-43ca-ada6-a04a7df1ce1a · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.651748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:38e7e813381ca48fa161c9cbf3668c2441bcd785dd65b32d7ae0ecb834777d57

Observation 5a86a0d0-18d2-4bec-9c6e-3b84c97c2383 · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:26:01.950624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:bbcbe37e8e733ce660d49ea30d3e6bd595bd60e990f49ba294d9905a21766c03

Observation 47cbdb7e-ffd7-4a73-95e5-b9f1ac60d283 · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Image quality assessment: from error visibility to structural similarity.IEEE transactions on image processing, 13(4):600–612

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.626496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:23ac343a34000768c4e4091c9974b1b49331bea4d5cbd20b83a330c10f325e8c

Observation edd76505-38a9-4b5d-8cb5-af12a9118c6f · outbound

This paper cites Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Open vision reasoner: Transferring linguistic cognitive behavior for visual reasoning

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.944505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:18bdb1bf056e357ca1a180eaebe44b5f9b050afe6a51c3dd280eff6d4c06250b

Observation cd8cd4ee-0ba8-46f9-be89-3e322b5230d6 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Depth anything: Unleashing the power of large-scale unlabeled data

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.661430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:dbc3558909cc28f28a81df59be2bc1a65ee8a43811bc61bad27b14e0e56cde43

Observation 4c6ad6c9-fdf3-43df-8778-46a1964bd261 · outbound

This paper cites Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Machine Mental Imagery: Empower Multimodal Reasoning with Latent Visual Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.935651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:18e38a91d48515dd21c49d8706da1fa761cadc0d3c03c60941052143c6fe1237

Observation ad7551bf-0755-407d-bbea-79925e110763 · outbound

This paper cites Introducing Visual Perception Token into Multimodal Large Language Model.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Introducing Visual Perception Token into Multimodal Large Language Model

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.904902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:48381841dd44eed793922b9549f970666a94bdc68cedb6f31c920ed20a7bcc33

Observation 4a740743-3438-4ae6-9ad0-ab9a1b75d28e · outbound

This paper cites Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Socratic Models: Composing Zero-Shot Multimodal Reasoning with Language

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:50:01.041344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:c61bd3f9106be7f5c617708ad22d45541605bfec9ccf9046c40affbf0c558e1a

Observation 65ead13e-0d5c-4b26-bcff-61e6f648efcc · outbound

This paper cites Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Improving the Reasoning of Multi-Image Grounding in MLLMs via Reinforcement Learning

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-11T10:26:01.917789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:28afd63e34a5e33e4d4eb471a2d0f4bdf4d464775c7bd9e8dbb905bf722c2f29

Observation e796cd22-198b-4d2f-bedd-16dcb0c9a8a4 · outbound

This paper cites Thyme: Think Beyond Images.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Thyme: Think Beyond Images

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:5f0d335ce01bf4404ef098e27bcb375d351022f45cce08985ef67a9ad177d31c

Observation 9d74e71e-3af3-426f-a6cd-fa668f1aaefb · outbound

This paper cites Vipact: Visual-perception enhancement via specialized vlm agent collaboration and tool-use.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Vipact: Visual-perception enhancement via specialized vlm agent collaboration and tool-use

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.958968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:2d0c50de48ac04f3488d94258270aee1d1ef76368d8647f48ed4cfaafc92ef8f

Observation 7c0852ea-11b5-4a98-be9c-18e7298cf73c · outbound

This paper cites Reinforced Visual Perception with Tools.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Reinforced Visual Perception with Tools

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:26:01.924607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:5b98dd28976284a310977fe5c89a4aaf849d031562e3326633594b0d8d58cf2c

Observation 827ec209-92dc-4c38-ae08-0b7fdcfc99bd · outbound

This paper cites Mainly, we provide samples of prompts for both frontier and open-source MLLMs.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Mainly, we provide samples of prompts for both frontier and open-source MLLMs

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.640749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:b84d7cfc9730f5046c5636fac99bb7923e5ee0e00b66f1d17af75566014e78b6

Observation d0ea9244-4b20-4433-b8aa-5651e641f6c2 · outbound

This paper cites left" or.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs left" or

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T23:25:29.644241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:9339b7d25095be9120f228b9909f4c0cad18d0a60646d137b2cba03c382c5cbe

Observation d94cf8ab-521a-4cb9-8777-7e17d42d456e · outbound

This paper cites 5.1 we discussed the quality of visual interpretation of current MLLMs.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs 5.1 we discussed the quality of visual interpretation of current MLLMs

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-05-17T23:25:29.676071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:7c6840b6d1a2b5539517f9242c52c5d914d76b9234e771ac4fd23421fe65f2e2

Observation 89665866-e434-43de-820b-3ef85228df5b · outbound

This paper cites an unresolved cited work.

Don't Show Pixels, Show Cues: Unlocking Visual Tool Reasoning in Language Models via Perception Programs Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-17T23:25:29.664380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T15:29:25.650175Z digest=sha256:d89f64f0b642431b688870fa3e9bd699f2cfc7cf7c9d46c08e5f5503d3cf22e2

Pith citing papers

No inbound Pith citation observations are available.