Pith. sign in

Paper Citation Record · LEDGER

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning

As of 20 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 0 inbound Pith citation observations for arXiv:2607.19790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19790 v1

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T11:47:29.180762Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

59 of 59 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved57
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 321e6055-520a-4719-b34d-b2d6c97c782f · outbound

This paper cites SPHINX: A synthetic environment for visual perception and reasoning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning SPHINX: A synthetic environment for visual perception and reasoning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:21.857897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:21.857897Z digest=sha256:f0c2d9b5f35321a246c351fb1fc53966efdea200b2f461d3e8e03d5941b15018

Observation 3666ff3d-ed4c-4ab9-9e4c-ec44845280b3 · outbound

This paper cites Qwen2.5-VL Technical Report.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Qwen2.5-VL Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.029634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.029634Z digest=sha256:e1c2dc28732f4c934dde8289819bc450581a41a9ade33608aabade6e3cd84c0e

Observation a17f5b4a-f1e3-436d-a374-4a8f01d3145e · outbound

This paper cites Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable Puzzles

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.196088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.196088Z digest=sha256:4d4b489e906c029733053bbb7dd4425d369e2cff52fffe5ef4277ca59ab077b6

Observation e831cd8b-5f64-4316-b892-d91a736961aa · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.325941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.325941Z digest=sha256:9de1917136fc1bff499b4f4543d1712df1527e75744636f9d890478f8107b2df

Observation a1fb70ef-eedd-4af2-8739-c04ccc6f6bc8 · outbound

This paper cites PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning PuzzleVQA: Diagnosing Multimodal Reasoning Challenges of Language Models with Abstract Visual Patterns

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.484880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.484880Z digest=sha256:d2b116f6c31378004b673e2a093883054d3607051332c56d6e7b2a8a2716f352

Observation e6451f18-8527-4441-b907-6fd29eee97aa · outbound

This paper cites EmbSpatial-Bench: Benchmarking spatial understanding for embodied tasks with large vision-language models.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning EmbSpatial-Bench: Benchmarking spatial understanding for embodied tasks with large vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.649103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.649103Z digest=sha256:9ba7e732a37aa0f9dbb3655b42c1ef755aa8ee8bcc90ea2ae18281b33c5ad977

Observation 53ec7949-035d-412e-a5af-11461813dd8d · outbound

This paper cites VLMEvalKit: An open-source toolkit for evaluating large multi-modality models.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning VLMEvalKit: An open-source toolkit for evaluating large multi-modality models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.815879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.815879Z digest=sha256:fc38496919ca1f142cc3a6cd6fa250f6d3ac693b404f951be52980c615a227b8

Observation 59bf5ae8-c8e4-4525-b8fe-c177fce0c22b · outbound

This paper cites VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning VisualSphinx: Large-Scale Synthetic Vision Logic Puzzles for RL

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:22.950972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:22.950972Z digest=sha256:d5b3bb464522e73b630dccc66a6ccbc2aa41c3a793733b76b93e89a7b73bd583

Observation fb24cba7-7b31-4c49-be5a-c6f5f1953e27 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.095760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.095760Z digest=sha256:d82683f2f5f7aae559bae5506906fb91378d0810d26927eab2070f94b091ab3e

Observation 2b0ad36e-671a-4e8c-94dd-4bb7fa6bfe25 · outbound

This paper cites Embodied reasoning question answer (ERQA) benchmark.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Embodied reasoning question answer (ERQA) benchmark

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.230236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.230236Z digest=sha256:e02f747f5a96217331deca359c39a2365ec02b6c2e678fc13dd614445b5e78d2

Observation 4f3530d0-3950-4c78-b670-57dbc90f4ffa · outbound

This paper cites Composition-grounded data synthesis for visual reasoning.arXiv preprint arXiv:2510.15040, 2025.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Composition-grounded data synthesis for visual reasoning.arXiv preprint arXiv:2510.15040, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.292929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.292929Z digest=sha256:16474be05b148928b3fb2ae6f3628750940cee33ea1a5a14d5730528f59abd65

Observation 141e940d-58ea-443e-9c67-19dc1834071c · outbound

This paper cites OpenThoughts: Data Recipes for Reasoning Models.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning OpenThoughts: Data Recipes for Reasoning Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.423108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.423108Z digest=sha256:86e2e67285bd7012285510495125f8ce0bbbadf28a98322da3e1e4b51b7b0f9e

Observation e7c1496b-f6de-44b3-b118-cf92ad9e5402 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.573664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.573664Z digest=sha256:08511d5e274e784a655fadf0b84f7af67e347479d158f6d20d2b555797fa0522

Observation 652dcc87-3c1a-4d41-9f59-0eb9898f3885 · outbound

This paper cites EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning EvoChart: A Benchmark and a Self-Training Approach Towards Real-World Chart Understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.720752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.720752Z digest=sha256:cc00f8c0dd8e1ce6bd6bc08917a41eec24763c2a3181f40b5558f2db404051d1

Observation 886d6178-534d-45c1-b65b-a4dc27699917 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.867456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.867456Z digest=sha256:546530ef520bb52bc9dae13d9217d4bb2273e5affdf29e1211610e7b893d17aa

Observation 6c6f66fd-0791-48b4-9449-dc25c1be387a · outbound

This paper cites Hudson and Christopher D.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Hudson and Christopher D

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.011421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.011421Z digest=sha256:a717bb999c00bd5f09ae8c4c49ee13391c22a5999be99659e361e3131af54b06

Observation 041995dd-775a-4187-87a0-be5a228ff0d4 · outbound

This paper cites Derpanis, Babak Taati, and Radek Grzeszczuk.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Derpanis, Babak Taati, and Radek Grzeszczuk

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.170473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.170473Z digest=sha256:1b832d0abbbbaf008f0fcf7b7f4efa72ca56ec28a9d054e4b7cffacb6922c514

Observation eace3887-7081-4ef5-a5fc-32b87bd679a1 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Lawrence Zitnick, and Ross Girshick

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.349175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.349175Z digest=sha256:d8441f02846a39ddb4bd069fecaabbbd66af2fa82766f511fdc5a6bb7f53eb82

Observation b36120e9-967c-433a-929e-a93860990c36 · outbound

This paper cites TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning TableVQA-Bench: A Visual Question Answering Benchmark on Multiple Table Domains

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.468658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.468658Z digest=sha256:fc48226df9c3ccf2a7ca7ccbb318117f5d42b82648001a876c88058eaa1d41d8

Observation 2cfc6284-3a9c-4a69-a9b6-5fefab123e99 · outbound

This paper cites Truth in the few: High-value data selection for efficient multi-modal reasoning.arXiv preprint arXiv:2506.04755, 2025.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Truth in the few: High-value data selection for efficient multi-modal reasoning.arXiv preprint arXiv:2506.04755, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.553141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.553141Z digest=sha256:a394c07d14097b74886bed6b5cf90d4d46bd5defbc860afcb3f2e910265843f1

Observation 5caa1d1d-90d8-4143-9426-0394621fd17e · outbound

This paper cites MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MoDoMoDo: Multi-Domain Data Mixtures for Multimodal LLM Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.668308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.668308Z digest=sha256:8512c5605593439acc2a4878472cb528cfef4698cf59ad72550b698c24de9321

Observation 28ff2ef6-090e-4df5-a6e5-eb57c1e34c4f · outbound

This paper cites Visual-RFT: Visual Reinforcement Fine-Tuning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Visual-RFT: Visual Reinforcement Fine-Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.777511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.777511Z digest=sha256:45e867f105d343e60c39d98822536e1345a98e46008113579a21c4d6bb47d312

Observation 72515d47-5a7a-45cf-b024-d808d159876e · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.880118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.880118Z digest=sha256:96f2c8d683f5dba857b33f32952415b42efdacf1dd7eac78d56ed7f3d245969d

Observation f9bbe934-2199-4c20-8eeb-c1dfb9c468a3 · outbound

This paper cites ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning ChartQAPro: A More Diverse and Challenging Benchmark for Chart Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:24.977601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:24.977601Z digest=sha256:a625eada386a3bb27213a7153a429d2f798c6286c3cf51474399e05cb29b5c1b

Observation a3b1d93c-aa97-4eb6-8738-e3ab016b5caf · outbound

This paper cites MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.109733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.109733Z digest=sha256:d974acb0039bc5aa4e12892cd5bae3a1c834710f8b1bb49633b8797312b66061

Observation 9b477d87-304a-4a48-8f9c-a964a55d4d88 · outbound

This paper cites Teaching CLIP to Count to Ten.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Teaching CLIP to Count to Ten

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.245565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.245565Z digest=sha256:9174101975b130c6760d92d4a6769f8da27ce1135719ccb4fedb30223b56629e

Observation b647749a-606c-4e14-a3e8-a0d72cfbb42b · outbound

This paper cites LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning LMM-R1: Empowering 3B LMMs with Strong Reasoning Abilities Through Two-Stage Rule-Based RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.396712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.396712Z digest=sha256:9e1e3a3bacb3d1a0d47405276b1234cc65e995c40487d99d2d9f561a81ad8c61

Observation 152ab05e-95eb-448d-8668-e4b5eabd59fd · outbound

This paper cites We-Math: Does your large multimodal model achieve human-like mathematical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning We-Math: Does your large multimodal model achieve human-like mathematical reasoning? InProceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 2025

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.546114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.546114Z digest=sha256:f226e34895a217911970561a6f283cd4d2f0a847d6eeedfd15660b0da6f60ee0

Observation 4402b3df-71fb-4359-b210-37b5950e1c81 · outbound

This paper cites Vero: An Open RL Recipe for General Visual Reasoning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Vero: An Open RL Recipe for General Visual Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.740386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.740386Z digest=sha256:8c7b956b7045a8ee81b06ffdd90af41e70b634e708e2ac9a105c2cb687ad201b

Observation 50a88d76-1de5-4863-adf5-8d7dbea322b6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:25.891265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:25.891265Z digest=sha256:dcb0bd63a97e2699b0b67ca3bc1486ceb6296d5d074c7ac04622925c39f13ec3

Observation 0e3ba484-1ff3-47ad-bb33-69b4695fd3bd · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.024291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.024291Z digest=sha256:a5f335aa367940b3b7eed620e74ef8694c3291f8066971111f9c754a31debf60

Observation ce460896-e6f5-4cb4-a343-8b5b2fdb2344 · outbound

This paper cites PhyX: Does Your Model Have the "Wits" for Physical Reasoning?.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning PhyX: Does Your Model Have the "Wits" for Physical Reasoning?

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.156753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.156753Z digest=sha256:ddf3815450fe4c6bfd1b596a98ba5e3948dc8a6abe59af8a9e8f41a8336c449d

Observation ca6f6a0e-2630-43d9-9ccf-325d95d6ef81 · outbound

This paper cites VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.293957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.293957Z digest=sha256:c7387b95a9cc28d5297c50c09a473e348043b11b53d2655369b97dc08c6a3897

Observation 7606f5b9-7f60-4e0d-9fcd-5addc7a25a30 · outbound

This paper cites Reasoning Gym: Reasoning environments for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2505.24760, 2025.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Reasoning Gym: Reasoning environments for reinforcement learning with verifiable rewards.arXiv preprint arXiv:2505.24760, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.427098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.427098Z digest=sha256:95b624063cdcab32b3b4a18b86355e1eb38ff9d3b850706df29c81a9122a2d07

Observation e6520cc1-8980-4761-8baa-842a387896f9 · outbound

This paper cites DeepVision-103K: A visually diverse, broad-coverage, and verifiable mathematical dataset for multimodal reasoning.arXiv preprint arXiv:2602.16742, 2026.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning DeepVision-103K: A visually diverse, broad-coverage, and verifiable mathematical dataset for multimodal reasoning.arXiv preprint arXiv:2602.16742, 2026

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.560618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.560618Z digest=sha256:85b3bf72a739ed98fbe75002f2fb99fe9bdea81b63902ac678f6851a7c3cfd28

Observation ec01bcf0-49c8-4923-bc8d-2df490edd88c · outbound

This paper cites CountQA: How Well Do MLLMs Count in the Wild?.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning CountQA: How Well Do MLLMs Count in the Wild?

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.652739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.652739Z digest=sha256:d20bd342a030ab018b5b1610f01cc78d9f2a5b016d86f0554a08343da9259161

Observation 4fbe2944-d714-4d0a-98aa-fc9554dcc30e · outbound

This paper cites Reason-RFT: Reinforcement fine-tuning for visual reasoning of vision language models.arXiv preprint arXiv:2503.20752, 2025.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Reason-RFT: Reinforcement fine-tuning for visual reasoning of vision language models.arXiv preprint arXiv:2503.20752, 2025

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.744210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.744210Z digest=sha256:c41af344306b29245a60b85de7ec0eeb559e7e968a5455d4c9d843135f58ee33

Observation e8d32674-ab7e-4ba7-b209-dad1ac92859d · outbound

This paper cites Game-RL: Synthesizing multimodal verifiable game data to boost VLMs’ general reasoning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Game-RL: Synthesizing multimodal verifiable game data to boost VLMs’ general reasoning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.848187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.848187Z digest=sha256:d494a27a6a8084a36d9099a1e3c300bfa85eeedb82428dfba5c41baeaf10bdd3

Observation 6d92f285-1413-4787-a8e2-f3df1ee96136 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.046319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.046319Z digest=sha256:d6f96d4173cd291f0f88a3c0b2f13e3a557f7a1171a4b388be226f7ce704d270

Observation 45a41f2a-3ca6-48b4-a9d4-0aecbea4f3c1 · outbound

This paper cites Traceable evidence enhanced visual grounded reasoning: Evaluation and methodology.arXiv preprint arXiv:2507.07999, 2026.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Traceable evidence enhanced visual grounded reasoning: Evaluation and methodology.arXiv preprint arXiv:2507.07999, 2026

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.159658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.159658Z digest=sha256:5e9a36d9be384f78065cf5240f8c6c480f399bb8a07141bcb38428de4327cb56

Observation fb550949-b86a-4754-b52e-ddb6301c20ae · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.252449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.252449Z digest=sha256:fe41d3be9bac76ce4d1823caadc94812ba3b13e12d75ed6735148ea11e82ff45

Observation 8986824e-c1f5-4f54-80b3-eae977c27c63 · outbound

This paper cites SpatialViz-Bench: A cognitively-grounded benchmark for diagnosing spatial visualization in MLLMs.arXiv preprint arXiv:2507.07610, 2026.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning SpatialViz-Bench: A cognitively-grounded benchmark for diagnosing spatial visualization in MLLMs.arXiv preprint arXiv:2507.07610, 2026

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.327542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.327542Z digest=sha256:e833a77d3c205d0fca0486e8f49145a34610e06e9a43b4ea3b6dbfe0c3efa7dc

Observation 5d28c6f6-f5b3-443a-99e4-67578dd105ef · outbound

This paper cites ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning ViCrit: A Verifiable Reinforcement Learning Proxy Task for Visual Perception in VLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.425487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.425487Z digest=sha256:8e06f1f5efbb2aa30b5c8a22e88e59ce89c20d1c7c43a71f761904df2ab3358f

Observation bf766f73-9c7e-47ba-8f1b-5d285e2986d9 · outbound

This paper cites SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning SoTA with Less: MCTS-Guided Sample Selection for Data-Efficient Visual Reasoning Self-Improvement

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.535628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.535628Z digest=sha256:880b576fde9a3ebe091f59dc80aa2519367287716ea0a146dc2a69b2a9e5889e

Observation 9442670b-fe8a-4e9a-af9b-e7838616f903 · outbound

This paper cites Blaschko.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Blaschko

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.697096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.697096Z digest=sha256:826f19c08bca926386a698a0f5c2c6373220375fdf508a5edaa73fac9adfd4b2

Observation b4bccc1e-36c1-44a6-a9ab-e0d217c2e645 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.812458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.812458Z digest=sha256:ce26fb77d6d246b185781f14e0dbca061f215b14d88627980a7166c74fc5986a

Observation d658f3bf-0e1d-456a-b0ec-cbf42d7578a0 · outbound

This paper cites SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.883967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.883967Z digest=sha256:7ee4a0618a5caca349ac3ad643137827064fe98345cc4f0388e41bb20d40133d

Observation 982adcee-b63c-4185-9dea-9853744bd561 · outbound

This paper cites RealWorldQA.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning RealWorldQA

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:27.981485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:27.981485Z digest=sha256:027fc8b5d6c7f1e2d268f527f83e27687eeaa9099f63a113248a8ecc855f6890

Observation 8200ef4e-1116-4813-bca9-cb034b6a2368 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.056919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.056919Z digest=sha256:1ce3d789a51291c1c9a1659c3c2123c538ea0d5c0d07bd30f75bd8f1e61798c4

Observation a38f450c-7479-419e-a93a-2d14d92ed709 · outbound

This paper cites WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning WeThink: Toward General-purpose Vision-Language Reasoning via Reinforcement Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.164827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.164827Z digest=sha256:848865f7a623c2952978cd0911c8e2847464d16297dd5270984cdb10fda6eead

Observation 0bcc56c5-31bd-4bf5-af0d-8d448254617c · outbound

This paper cites TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning TRON: Targeted Rule-Verifiable Online Environments for Visual Reasoning RL

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.286163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.286163Z digest=sha256:f95cb98a92a1c39fbd71ee5aa76e6ef9e42dbfceec1ae9203d24061dd3d77ea9

Observation 06628b5c-4a62-4556-80de-50f448c77ed3 · outbound

This paper cites R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning R1-Onevision: Advancing Generalized Multimodal Reasoning through Cross-Modal Formalization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.392380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.392380Z digest=sha256:afa4a8eaeeec28f9c14fbd4bd81b23a76a4f8a8a82cb5a7aa3a569ffbfb5cca7

Observation 12823ab1-fbfe-4e24-b384-51cd57eeed34 · outbound

This paper cites WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning WorldBench: A Challenging and Visually Diverse Multimodal Reasoning Benchmark

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.553275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.553275Z digest=sha256:5e188b3d083f71afdc693b7862de4257f29b9ba5a300be45d3cf17561e2870f4

Observation e9db77c4-c4b4-4367-b9ce-fa6413821832 · outbound

This paper cites MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MME-Reasoning: A Comprehensive Benchmark for Logical Reasoning in MLLMs

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.673571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.673571Z digest=sha256:b3a87b18c839139deb429e0f0e24fd7469f9ddb54bfd0c625f964e91198de1c6

Observation 18c75216-e092-4c20-b998-cb3ea5b7b8a5 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:28.785230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:28.785230Z digest=sha256:8c328fd87f26481ef87600c89f4dd6e5519f7b4b1e35971621399059cd47b9ba

Observation 8bb04209-2a80-47a3-8067-4b8e91a10313 · outbound

This paper cites Xing, and Zhiting Hu.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Xing, and Zhiting Hu

Reference 56

Resolution
verified exact
doi, observed 2026-08-01T11:48:19.392604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-01T11:47:28.924987Z digest=sha256:e36f0563b1b7643248f3abd7e812bbb24a0c58aab219a8d224aa91dc79e7be52

Observation 1fff77ec-b775-4dbf-bb49-d764082b07b7 · outbound

This paper cites Task Me Anything.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Task Me Anything

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:29.045426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:29.045426Z digest=sha256:f4e0dcab366564e57a98cc7d0f39a1a0fff969404f52770a11d6b286b8b4498d

Observation 7f8ef853-976f-44fb-9605-cccd4b3d12ff · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 58

Resolution
malformed identifier
no resolver link, observed 2026-08-01T11:47:29.180762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:29.180762Z digest=sha256:2a723eb18684fc1f6154e82243f1de90f8e4fa902c9de90d32dcfb54591fae50

Observation 8da32b69-1d74-4955-b2a0-d1a1387db81e · outbound

This paper cites an unresolved cited work.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning Unresolved cited work

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:26.938715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:26.938715Z digest=sha256:e54b7c13ef5e3c8ab9d7ee3ef85371acf3cae0306a5e0892c7def9a2a4e1458a

Pith citing papers

No inbound Pith citation observations are available.