Pith. sign in

Paper Citation Record · LEDGER

Perceptual Flow Network for Visually Grounded Reasoning

As of 7 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2605.02730.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.02730 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T18:40:55.753827Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact36
  • verified fuzzy29
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d03f9924-7868-4d02-aa94-fef5ffe3dcd3 · outbound

This paper cites Qwen3-VL Technical Report.

Perceptual Flow Network for Visually Grounded Reasoning Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.808532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:e56c1c66cae47a2a9cd20f8df00ba4fa3a1535f57f9b405612cf9933b37e8ac2

Observation 51ace5b3-3943-44d3-b41b-1ba05d653159 · outbound

This paper cites Qwen2.5-VL Technical Report.

Perceptual Flow Network for Visually Grounded Reasoning Qwen2.5-VL Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.858186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:dc696e528b582df1ead8341bfdd3e3b1e9a2999766f362e2c6bf7de1c532e2d2

Observation a1f5fec5-32a8-40d5-9643-6a4e870645a7 · outbound

This paper cites Vicinal risk minimization.Advances in neural information processing systems, 13.

Perceptual Flow Network for Visually Grounded Reasoning Vicinal risk minimization.Advances in neural information processing systems, 13

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.330589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:7e44125ea032cf81963d1003616fc2d0fab15f06aa07b172fcbffa69898bff75

Observation 690df2cb-1b5e-48c5-aac0-1f8201173ee9 · outbound

This paper cites Multi-object hallucination in vision language models.Advances in Neural Information Processing Systems, 37: 44393–44418.

Perceptual Flow Network for Visually Grounded Reasoning Multi-object hallucination in vision language models.Advances in Neural Information Processing Systems, 37: 44393–44418

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.326972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:de65c547de2577c2c6f21d1e14afcb427c1c5100f337df895e58da6fec816146

Observation 4ccc9680-77e0-4b08-ba81-be07cd2bc7e1 · outbound

This paper cites Seeclick: Harnessing gui grounding for advanced visual gui agents.

Perceptual Flow Network for Visually Grounded Reasoning Seeclick: Harnessing gui grounding for advanced visual gui agents

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.322948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:1eefb8365207df89b5af8d10c06e8b26b530ac3829f00e9ba68ebb3650206726

Observation 6c30ec8a-d0a3-4433-9c08-8b7983b501fc · outbound

This paper cites Gemini-3-flash.https://deepmind.google/models/gemini/flash/.

Perceptual Flow Network for Visually Grounded Reasoning Gemini-3-flash.https://deepmind.google/models/gemini/flash/

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.334009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:e7c5ca26ff415b1218b9f29e86554794b45e6031706d7d05060ab6bf3e219979

Observation 5255d94d-c2b1-4402-bafb-eb843a23a91b · outbound

This paper cites Gemini-3-pro.https://deepmind.google/models/gemini/pro/.

Perceptual Flow Network for Visually Grounded Reasoning Gemini-3-pro.https://deepmind.google/models/gemini/pro/

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.319331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:cf3fb351575363e575f72ab9c7db354e9c864c695bf2ba8d866244123170766a

Observation d66df7ea-1683-4d03-acc4-3baf5c1e37cb · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

Perceptual Flow Network for Visually Grounded Reasoning Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:09:01.446022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:3f8ce87e68f64a48505e89d16ae7a0b98774b257d768cfae40b64cc731020e0d

Observation 92d355c1-b648-4ea0-9225-e99b3c3af553 · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

Perceptual Flow Network for Visually Grounded Reasoning Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.375158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:ac3a087082bcb1da65218a0caafba152d0a571685d92a0fdcb4c7ed1ac93d93f

Observation 42d13a26-4565-415f-a9bd-fbdba5ef2483 · outbound

This paper cites Detecting and preventing hallucinations in large vision language models.

Perceptual Flow Network for Visually Grounded Reasoning Detecting and preventing hallucinations in large vision language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.366062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:9bd9bda01da502813a2e7f30c77e11619e6f846b0132aeb08246c9cc7a8a051e

Observation 00b472c5-175a-4f87-a9a3-e5c28814cf46 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Perceptual Flow Network for Visually Grounded Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:15:38.975491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:c4aef692b1e594ad6330fd04df459fbda752dcbfa1d7d69753f448c0e0fabbb4

Observation 3853f9ff-e339-4bf5-a7cd-9cf5f15b5235 · outbound

This paper cites Seed1.5-VL Technical Report.

Perceptual Flow Network for Visually Grounded Reasoning Seed1.5-VL Technical Report

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:26:06.820639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:53afd98bd2d337e525f8df79e76d3d2fc0cf44b3a0612f57225f5527a8c0ff78

Observation 9bae1b68-dbbf-4989-ab98-e80bb6049a51 · outbound

This paper cites DeepEyesV2: Toward Agentic Multimodal Model.

Perceptual Flow Network for Visually Grounded Reasoning DeepEyesV2: Toward Agentic Multimodal Model

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:32:29.694372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:b1c095445d711d1100ab0084b93983a053a314b3dbfa81480015454ee0ebc523

Observation 8ece9cdd-0f46-4117-b86e-b21a0f0e7bd6 · outbound

This paper cites The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use.

Perceptual Flow Network for Visually Grounded Reasoning The Dawn of GUI Agent: A Preliminary Case Study with Claude 3.5 Computer Use

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.873140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:48cd3dada78433e972d03641ad5cecec807e40b9930cbcabd470901c25418a4c

Observation 6a455e61-a13b-409e-a01d-626a6a0803e4 · outbound

This paper cites Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379.

Perceptual Flow Network for Visually Grounded Reasoning Visual sketchpad: Sketching as a visual chain of thought for multimodal language models.Advances in Neural Information Processing Systems, 37:139348–139379

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.362470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:82dd8f3d30c2df73278cfad9d943b0aace4292967c64953500a33cfd2c6756c7

Observation 8243bba0-66b1-4c28-95a8-0c40596e4190 · outbound

This paper cites VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought.

Perceptual Flow Network for Visually Grounded Reasoning VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.910256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:d679cf3ea730263d33de880bd431085d39d1620992906bc78735fa985753c6b0

Observation 43424181-65a9-4fe2-a15d-3016eff905ca · outbound

This paper cites A diagram is worth a dozen images.

Perceptual Flow Network for Visually Grounded Reasoning A diagram is worth a dozen images

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.369220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:766e222d1480af4973147c099eed1bbdecbbca5e0c8d2434f5021b4efc176804

Observation b34f3014-427a-45ac-8044-9e2b38503b96 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Perceptual Flow Network for Visually Grounded Reasoning Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.372209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:bb7a85039c4cba71293b66d735cc455d5bbaa16a127a93dc3c0b2a76a229dde8

Observation 82811d9a-860a-4dd9-a5ab-05aed6e62bd3 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Perceptual Flow Network for Visually Grounded Reasoning LLaVA-OneVision: Easy Visual Task Transfer

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.838668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:923ed539fa961c8c53783c9b5f73f946555c8f228614119c1a4d762294211123

Observation 06c94115-6e53-466d-a423-fc62d6c0f880 · outbound

This paper cites Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding.

Perceptual Flow Network for Visually Grounded Reasoning Dyfo: A training-free dynamic focus visual search for enhancing lmms in fine-grained visual understanding

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.382120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:6e38dde4d77c0ebfdb3a74f07f93df26ec85cb81b747c9c46b1f89a618ab0f51

Observation fbd15183-87e5-4c1f-90e4-479d0adc0796 · outbound

This paper cites Screenspot-pro: Gui grounding for professional high-resolution computer use.

Perceptual Flow Network for Visually Grounded Reasoning Screenspot-pro: Gui grounding for professional high-resolution computer use

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.350037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:8d49bdf5f4feb9aacee20e9da3eebe64cae324801649c440296e53b3de4ac138

Observation 1142182e-cf02-4766-a44b-f99b9fedb515 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

Perceptual Flow Network for Visually Grounded Reasoning Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.923328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:cf05f66dc55bc0911063494158c94f4240a2880767ea35b2aa8d0153ae837767

Observation 60f25f82-1032-411d-824b-b4bfac8cc294 · outbound

This paper cites arXiv preprint arXiv:2603.03857 (2026).

Perceptual Flow Network for Visually Grounded Reasoning arXiv preprint arXiv:2603.03857 (2026)

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.953868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:16b8eb943e3b82ba5bfd2b520f6c9f86b88a7a76218897be8b578b6c2d7ebe8a

Observation 9bbe7f85-3904-4c30-8f5b-24271e8800f0 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Perceptual Flow Network for Visually Grounded Reasoning Evaluating Object Hallucination in Large Vision-Language Models

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:44:10.072171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:8d80e889afde51acb8c88676cbfa04e96f0293b496e060369f40b629d0d6df5b

Observation 1e8d896b-5d56-4cf9-a435-ea9827dbe3b1 · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

Perceptual Flow Network for Visually Grounded Reasoning A Survey on Hallucination in Large Vision-Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:10:10.379182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:8f0a86782e795e3d9a5ecf8caffbeac72aaae28ff45fcfeedc34655e0b57cf2e

Observation 8247fb7d-0ad2-44dd-b76f-cb60b85bacab · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916.

Perceptual Flow Network for Visually Grounded Reasoning Visual instruction tuning.Advances in neural information processing systems, 36:34892–34916

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.342731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:051fe931915a5febfbaf4b9761ad10bf2b1e662688811f5bbfd4299df2779757

Observation 495c17ca-aa93-4950-980f-f50e411c6f5f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/2024-01-30-llava-next/.

Perceptual Flow Network for Visually Grounded Reasoning Llava-next: Improved reasoning, ocr, and world knowledge.https://llava-vl.github.io/blog/2024-01-30-llava-next/

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.338815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:d9f2c04ab711fadd2ed421cb4d3bcf0aa1b05af0f70cf4ad7a2ad0ea0d859033

Observation 280ccb43-1365-4cec-8d79-965f4e89551c · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Perceptual Flow Network for Visually Grounded Reasoning Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.346600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:280bfb9a4f1fc2a80e7cbe3d19ae8d206235531e3faa0ab0106e8a6b8536c084

Observation 1392a791-a650-4c04-9439-6a34af18eab3 · outbound

This paper cites Look as you think: Unifying reasoning and visual evidence attribution for verifiable document rag via reinforcement learning.

Perceptual Flow Network for Visually Grounded Reasoning Look as you think: Unifying reasoning and visual evidence attribution for verifiable document rag via reinforcement learning

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.939789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:6d99470c64b3df5891da71a06504d416249a748682a078177d685eca27c5ed5f

Observation 963ceed0-7165-4296-a033-ba50d2872694 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233.

Perceptual Flow Network for Visually Grounded Reasoning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pages 216–233

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.358600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:19054b8980387bad2c024d007d65f7c30bcedee5937299bf94afc8fff4542ad3

Observation db48f923-6577-4fbe-abaa-69289c49572b · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Perceptual Flow Network for Visually Grounded Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T22:16:16.677111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:d8f9e180c5c5392e24c030f984efeafcbfe5507888c06b661b8e55fd4c3eb93f

Observation 6c77e42a-913e-4405-98e4-2bfc083b8fdb · outbound

This paper cites Visual Agentic Reinforcement Fine-Tuning.

Perceptual Flow Network for Visually Grounded Reasoning Visual Agentic Reinforcement Fine-Tuning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.959109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:f8624b5310743ba2beb86b24db3710260c741007d6687e25eff98af1c13be770

Observation 9f4446c5-3ced-4497-86c2-aa4194159e55 · outbound

This paper cites Decoupled Weight Decay Regularization.

Perceptual Flow Network for Visually Grounded Reasoning Decoupled Weight Decay Regularization

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:15:38.862453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:5ee8cb9d8cea6d96f7128c7514aca8863b645a4d839f959df7a3be503c556cfd

Observation 34050e6a-4513-47ae-b94f-3533ad73f104 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

Perceptual Flow Network for Visually Grounded Reasoning GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:10:58.174913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:0fff027463fa73890471864ddf79d18d9c2f29642560a30d9fa8a2b2178df38c

Observation b3b68f45-4d24-4d24-9f03-b9ded88ddce3 · outbound

This paper cites Struvis: Enhancing reasoning-based text-to-image generation via thinking with structured vision.

Perceptual Flow Network for Visually Grounded Reasoning Struvis: Enhancing reasoning-based text-to-image generation via thinking with structured vision

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.891292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:3f115395ea61434c1329ce67efa6e19b04a9f78e7e4d9831df3779c5fbe2cbda

Observation 42b92253-1357-44fc-b7ea-c07a2e37b357 · outbound

This paper cites Learning gflownets from partial episodes for improved convergence and stability.

Perceptual Flow Network for Visually Grounded Reasoning Learning gflownets from partial episodes for improved convergence and stability

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.328289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:989ab318e985acb4747932d520ae9ba263206b795aaabc2ba19f91b4893b8f3d

Observation 5cc5c6d5-5abf-424e-a0bc-7ba0cab2f65f · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

Perceptual Flow Network for Visually Grounded Reasoning Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.332057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:8b21214088e2a26167601bf6e8c11b8bef69e3c062f7d5e3abbfb808b8c0ba30

Observation 354a0b71-94b0-4996-973c-153c98fbb277 · outbound

This paper cites Openai-gpt-4o.https://openai.com/index/gpt-4o-system-card/.

Perceptual Flow Network for Visually Grounded Reasoning Openai-gpt-4o.https://openai.com/index/gpt-4o-system-card/

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.321161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:b923a5615c23492b42d622867135f5d9cf1203ed844b0f814c429b73b0640b5b

Observation b3291600-312c-433b-99c8-fd5b7f602867 · outbound

This paper cites Openai-o3.https://openai.com/index/introducing-o3-and-o4-mini/.

Perceptual Flow Network for Visually Grounded Reasoning Openai-o3.https://openai.com/index/introducing-o3-and-o4-mini/

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.317564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:23c7c0104c17be70b0a0b994a98936d23a5908bdbfa7656f042732f46ec74669

Observation 92d9c17e-9265-407d-95f0-8749e883c5c7 · outbound

This paper cites Operator: A computer-using agent.https://openai.com/index/operator-system-card/.

Perceptual Flow Network for Visually Grounded Reasoning Operator: A computer-using agent.https://openai.com/index/operator-system-card/

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.324686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:3fee700048cbe950805676b10e4ad15264fcb4a4b911adea85421fa0cd4c3ce5

Observation 92fc9604-0070-4883-bb9c-4ce08f170083 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

Perceptual Flow Network for Visually Grounded Reasoning UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.931163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:4fba7be8f989cdffb0df2a2da68d9e4f1b563a845b9c7019d8abc0a5ec0518be

Observation 1f299891-0086-45c3-9564-655ca1528a8c · outbound

This paper cites Learning transferable visual models from natural language supervision.

Perceptual Flow Network for Visually Grounded Reasoning Learning transferable visual models from natural language supervision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.335435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:0da0afd0c1a166ce9f04f31da84c7a6e6d68799f6d4e196b4089112bbf9e4602

Observation 68831d3e-7207-4926-9b9a-ce8a33a8ec54 · outbound

This paper cites Grounded Reinforcement Learning for Visual Reasoning.

Perceptual Flow Network for Visually Grounded Reasoning Grounded Reinforcement Learning for Visual Reasoning

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.942584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:3a77b43c92c0cb8d414fe19e83988c54db015a54616882cca3c53aebf9634c1e

Observation a81389c6-a2b6-40a6-b5d9-93a8301ae4f9 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Perceptual Flow Network for Visually Grounded Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:13:57.602083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:cf994bf63daa495954e36b9398214b1866ff3e14d8d2521736f18497aa31ce3b

Observation 50515f0d-50f5-4fa3-9cf0-7db2b9113c5f · outbound

This paper cites Codedance: A dynamic tool-integrated mllm for executable visual reasoning.

Perceptual Flow Network for Visually Grounded Reasoning Codedance: A dynamic tool-integrated mllm for executable visual reasoning

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.927047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:6707d3a895a24ffd47b578b85cee70cb7b343ba5012c10cf81521cd4d7e6f9df

Observation 51eaa465-a263-429f-ad8b-f108bfaeee45 · outbound

This paper cites Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning.

Perceptual Flow Network for Visually Grounded Reasoning Pixel Reasoner: Incentivizing Pixel-Space Reasoning with Curiosity-Driven Reinforcement Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T02:22:27.090901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:1c6a5c418dbb38ce48dbfd013d3b18316a48cf5ea200ab57093e7e5fce84f460

Observation 143290ce-9240-42f3-8932-36e849ee8ca6 · outbound

This paper cites Kimi-VL Technical Report.

Perceptual Flow Network for Visually Grounded Reasoning Kimi-VL Technical Report

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:06:31.748869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:966bdb1beb55559aa79e68270cc900fed3fae26a1a973eee6a5cf9b31f5a342c

Observation 5b11a51d-7e19-4344-82c4-d10c302e3b24 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356.

Perceptual Flow Network for Visually Grounded Reasoning Cambrian-1: A fully open, vision-centric exploration of multimodal llms.Advances in Neural Information Processing Systems, 37:87310–87356

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.378544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:77c83519f3e3d1c3dbbe5948b2625ed32ad9164c39a38454f3ac4462725dae8b

Observation b4da331a-8359-4c22-90a7-2aea09463702 · outbound

This paper cites Attention is all you need.Advances in neural information processing systems, 30.

Perceptual Flow Network for Visually Grounded Reasoning Attention is all you need.Advances in neural information processing systems, 30

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.313927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:5ba1e37679a555510bfc30f2bbbfe2d26b0d5874ba404f28040feebcb9f07b47

Observation ac07db42-e67a-449d-9b37-61c9dd0d785d · outbound

This paper cites Trl: Transformer reinforcement learning.

Perceptual Flow Network for Visually Grounded Reasoning Trl: Transformer reinforcement learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.353364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:e687d398b8d50b97fa0c476609ec28f6df789d75c8a578536a338914dc65e72c

Observation 23e3e54a-db4d-488f-aa32-3c83635bb105 · outbound

This paper cites Traceable evidence enhanced visual grounded reasoning: Evaluation and methodology.

Perceptual Flow Network for Visually Grounded Reasoning Traceable evidence enhanced visual grounded reasoning: Evaluation and methodology

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.824685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:858e9c1da493df6c4b7ed683c81133f47972b2873b4685247f41b1bdf784b68c

Observation 9f5d6f7c-2247-4503-9983-1672ef99f06f · outbound

This paper cites VGR: Visual Grounded Reasoning.

Perceptual Flow Network for Visually Grounded Reasoning VGR: Visual Grounded Reasoning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-09T06:15:38.970783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:bf616cd17db45a09ef3b3a129b2526aee629d9c1f23f9e52d7d0ec50c0ea4781

Observation 9fd20e8d-c213-49f1-8e2b-e14185c83faf · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169.

Perceptual Flow Network for Visually Grounded Reasoning Measuring multimodal mathematical reasoning with math-vision dataset.Advances in Neural Information Processing Systems, 37:95095–95169

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.310577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:48305f51ecb0e9fc0aca0ab7cf7eaee621b451d8d442e33ee1e7bb793f4ade80

Observation 64910dce-dc69-4630-a579-ba2f2aeebb41 · outbound

This paper cites Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models.

Perceptual Flow Network for Visually Grounded Reasoning Divide, conquer and combine: A training-free framework for high-resolution image perception in multimodal large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.306716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:b0c1153ae683fe85f2f711718bbfe0d743bc4f8a0ce37fc34e922a36559d7bc3

Observation df048a35-a7fe-4496-b6f0-40e77bff5d6d · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Perceptual Flow Network for Visually Grounded Reasoning SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.894927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:e3791f18ae77d85f13863d5f97b0ecfb870d270da19b5af54f742e4a501acaf5

Observation 11cc80b7-b227-4e7b-9ad3-2ec7c91d7fd0 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

Perceptual Flow Network for Visually Grounded Reasoning V*: Guided visual search as a core mechanism in multimodal llms

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.303415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:f70dc5dd79dc896330d6546e1656a37408cd75335144297f742ee69e0f1ecf98

Observation 42bb3cab-8a93-4d32-b092-b129026b36e0 · outbound

This paper cites Os-atlas: Foundation action model for generalist gui agents.

Perceptual Flow Network for Visually Grounded Reasoning Os-atlas: Foundation action model for generalist gui agents

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T04:17:21.299438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:fd6f9f69af56e57c634f2342a3c56895240e6f35a62bfceb1573fdfed288b8fe

Observation 927cc72e-da19-4c93-bf20-e80bbb8f0d99 · outbound

This paper cites Vacot: Rethinking visual data augmentation with vlms.

Perceptual Flow Network for Visually Grounded Reasoning Vacot: Rethinking visual data augmentation with vlms

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.851284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:dba776d2c141bc7469aff0299d3e744712db78e18ecb53d27f67459bf769e2a3

Observation 9868b5db-7fe2-451f-b3ad-670da9d701f7 · outbound

This paper cites Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement.

Perceptual Flow Network for Visually Grounded Reasoning Zoom-Refine: Boosting High-Resolution Multimodal Understanding via Localized Zoom and Self-Refinement

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.906515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:59b739d8ca312ffc32aede0e1aa40e4c6fe2fa7e13e41f3fe096f0655cd2708b

Observation 6fc1e1c5-a631-47c3-9351-f136557fccc2 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Perceptual Flow Network for Visually Grounded Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.794677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:32b7ce9f18ad208128b358f4c27263290116be19d80f5e580075a7dbf28b1e4e

Observation c9983571-068c-4d3a-acbb-971c8aff9e13 · outbound

This paper cites MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?.

Perceptual Flow Network for Visually Grounded Reasoning MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans?

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.958879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:475eef29661ea27173bec3bb98106b1bcf55fd240238c4c2153dca98922b50af

Observation e91595cd-287b-4832-81ae-129ed0ae54cd · outbound

This paper cites Thyme: Think Beyond Images.

Perceptual Flow Network for Visually Grounded Reasoning Thyme: Think Beyond Images

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:33:29.495992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:4d2a696ff914366aa9e65dfb1e4c7b38ae95aeb8b6c089067e0fc686d4dde6bc

Observation 4093ded3-4e22-4222-a32c-f11ff7e797cf · outbound

This paper cites Mirg-rl: Multi-image reasoning and grounding with reinforcement learning.

Perceptual Flow Network for Visually Grounded Reasoning Mirg-rl: Multi-image reasoning and grounding with reinforcement learning

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:15:38.800143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:ab5d57142e956d2ec78ff1d73dd0e5963e4ac7c8e809dc0c43864c03b90ae9c0

Observation 8ce60f44-4638-442c-a7ab-c7acf4714fda · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Perceptual Flow Network for Visually Grounded Reasoning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:40:02.344422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:9b7ba64fd0fa42eb9076b81ea8019985b8b989db924c255721ffed6705503905

Observation 4bb3097a-c652-4277-8c5a-2e34a4f2b928 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Perceptual Flow Network for Visually Grounded Reasoning DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:42:57.181531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:2f42d204707020ad600b5bb7aeb2d52204cac386ec5810dd77f8da09659f912f

Observation 4b57d5ef-a71a-4ec7-b020-76f0ec0581ef · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Perceptual Flow Network for Visually Grounded Reasoning InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:41:08.498970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T18:40:55.753827Z digest=sha256:74cf7ef63b8a216c1d34b84cfb390311502042b9e2f8794c36d8b1852129d660

Pith citing papers

No inbound Pith citation observations are available.