Pith. sign in

Paper Citation Record · LEDGER

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

As of 22 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 8 inbound Pith citation observations for arXiv:2506.07235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07235 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:45:30.735139Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:31:41.386191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:43:28.625652Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26949dd0-946c-4b20-8aae-73a130af3245 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Blink: Multimodal large language models can see but not perceive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.201843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.587797Z digest=sha256:04b67e2002a90b66445d125e2812fe8c4f8b2e6d40eda5b61a6c6644e8785a2b

Observation 252017fe-5670-4cb6-b23d-e17e2c40e8b6 · outbound

This paper cites V∗: Guided visual search as a core mechanism in multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V∗: Guided visual search as a core mechanism in multimodal llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.188988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.591770Z digest=sha256:d8d96d7a98ab7182cf879431e3c7002d264280c6d56ff8123a1cf2c6f41ab369

Observation e674edb9-6e54-4a7f-aa85-33d81f32411a · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.595949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.595949Z digest=sha256:bdd33ec156990ebbd7e5990a3a07787ed19de1bfa1f9ea80f6f25a799f59dc6b

Observation e5b51c26-c286-46df-9016-544557ae8b24 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.600959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.600959Z digest=sha256:dc2be016e4f517d2521af3f0aed86e4cc3953afedcb46dade1ecf27a041263f5

Observation bbb19b7a-cbac-4017-9c6e-430cd949d879 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.605338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.605338Z digest=sha256:d8ab9ac191177e5640841d4ca21a752e863554f17636cbec30ca9bc336967658

Observation 52e12acf-5b2a-4255-8ca5-94afd5e420d3 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.610059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.610059Z digest=sha256:220af059ffd75c308e1959dd7ca260c2a0da461a19a64bfd12b4b7af90cdd12a

Observation 5d8a4253-37a2-4eec-abcf-5c91400e7576 · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.168232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.615181Z digest=sha256:a76a434cf0f286f77d93debd58d6ed550144f6f80903d51b249f68a7c057139a

Observation cbcf874c-d05b-42ec-a1f6-e7368d67ad0b · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual programming: Compositional visual reasoning without training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.619311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.619311Z digest=sha256:01988d39af360568ed69273a2c29eeabae0ffef3e7c068f17c12e57de859d8cf

Observation 9b9772bb-683c-4563-8847-1e8b248f6754 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Vipergpt: Visual inference via python execution for reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.146028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.623036Z digest=sha256:c68d5751396d9d3d3743369a1935e663f54942c4b3929e9a5d58ff1e37cd79d8

Observation 92032de7-d6d7-48c5-9a3f-a5cf9317df0b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.627374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.627374Z digest=sha256:d342853bd4cd1191b6809197e9fd23d6602f2fcfa69db4f00cbe25c5c0a82bcd

Observation b523a58f-2eef-46ee-b336-3950e403d8e8 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.631557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.631557Z digest=sha256:158e4f9dfe2335fadbe21ffe8004db8b7cd2a9290979b0b39fee6382856e48d4

Observation bf3df506-9ebc-4103-a789-956e840cebd0 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.635922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.635922Z digest=sha256:861867615f9a1facf218c09ac5bee5e2bc140b070dfc72ceadb211156cd5c253

Observation dcf6d6da-1844-4d90-a1f5-fa4898a2e6f5 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.640388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.640388Z digest=sha256:0fdcf6539a26be3f44a76d8b8d20c2950fd6d28e3438367fe445da788887b73a

Observation fc168fb3-3060-4879-b51d-a94c1a8661d8 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Making language models better reasoners with step-aware verifier

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.644103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.644103Z digest=sha256:8bc5421f9c5390f051b37bda66b751d2f1d6ce1295805cfe92f7004461576fa7

Observation cb1f1667-6ae4-482e-a7ad-06dfc7a1262a · outbound

This paper cites Let’s verify step by step.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let’s verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.648369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.648369Z digest=sha256:194c4601f6da82c42575ac13f5516d2ef2501cf2abc5a13d17e9c3b661f4a556

Observation 3e4fe20d-b57c-481a-b5cd-c493fd0e6d42 · outbound

This paper cites Solving math word problems via cooperative reasoning induced language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Solving math word problems via cooperative reasoning induced language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.109578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.652126Z digest=sha256:e9eed15610f07fb6ff2cb7f67db7444147fbc94bb0b5e43cd873c456388bc19e

Observation fb622090-c761-4f60-8374-09194f456762 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.655336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.655336Z digest=sha256:ff38b1454c55b2f041192d42f0ee8c723d4176b3cd931804177b31fa9c68f542

Observation 2e143f65-c3c8-497e-892e-f6efc4614f71 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.658831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.658831Z digest=sha256:3e47acedfb32e0dcd765fd5fabc94dae4d5ff00dae34f23df4554edabb279dac

Observation 68ac5975-e95e-4064-b93e-b95e536655c7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Fine-Tuning Language Models from Human Preferences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.662617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.662617Z digest=sha256:8ee04d4630d4675aa153af356e0ff3d18b6c4f4e3832f0add630ef4f061ba2f8

Observation 184ebd8e-b7e4-47db-80de-756321b7cf7b · outbound

This paper cites Rlcd: Reinforcement learning from contrastive distillation for lm alignment.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Rlcd: Reinforcement learning from contrastive distillation for lm alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.097704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.666869Z digest=sha256:688d4013ee2617bb43b920410afe2eccf11317888b59bad2d9349eb4fc3de6a9

Observation 3ba739d5-16d2-4232-bc94-8f028c22e319 · outbound

This paper cites Pretraining language models with human preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Pretraining language models with human preferences

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.084696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.671380Z digest=sha256:93449939101646b79929ee652a23c6aadd9edb81b27b9baadd9e952bf1c07fa0

Observation cdee5b0e-b45c-4eef-b8b3-e02f34ede260 · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.675404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.675404Z digest=sha256:ef4f8eb41d017ace9bc2c8ba346af6deac70f51873077639f098ce33cfd30940

Observation b221ca03-9306-44e2-b182-2b87ba767138 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.679605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.679605Z digest=sha256:de70b65df2a3e41f6b1c851fca0d6d04929387ac0e18cef98e3ae21c947d2808

Observation 0e2ad5c0-c823-4ec9-b278-869e749f1dc5 · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.683409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.683409Z digest=sha256:cf674ea690cfc2a698c023b3f7961b538cebae6bd2f523393c24eb674fbae6cc

Observation 4bde0720-d8c2-4b59-923e-cfb12b1e46bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.687744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.687744Z digest=sha256:339e6ebbff13cb7e40a88a9aef4e109fcb7f4ee74242fdec705006df2d658f98

Observation 7a1641a5-f9cd-41d4-b1f7-03d615ea423b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.691615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.691615Z digest=sha256:41da15868a0965e4f0a5303af1434c5f2573fc5479dcf9ceb8da07b21e94d6a2

Observation 5afb27e4-9bc5-454f-a275-f37192943a9f · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2.5-VL Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.694953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.694953Z digest=sha256:93dadd05520d0e1f4d334f077e82a7d552f11a4885202cd782010312aee657f2

Observation f957249a-e9a8-470d-837c-08b495156d34 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.698822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.698822Z digest=sha256:dea7561286233e90c3cf05b31df9acc18de5e20f8422cd38b1598964f9f80510

Observation 459c0db6-b89b-4b85-948e-7185cb7681f2 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.702326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.702326Z digest=sha256:2f767801cf803049539a6835f6d7882cd8f7c40bdf38f630df6f41191f1ebaad

Observation 65f1e13d-11f2-4578-9337-f6ffec9cf394 · outbound

This paper cites MMFactory: A Universal Solution Search Engine for Vision-Language Tasks.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.706182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.706182Z digest=sha256:585d47fc8c21bd965617d7a066fc503ff0c430617f77bd3ce08dd74662df0793

Observation 0d54322b-c595-4edf-bbea-a1c08399fd84 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.709569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.709569Z digest=sha256:4c7aab3010d4bc366f954579b40107cea4f83ddf07cdd4993ecb4c3ee6166d2d

Observation 2c26b258-5e7b-4a1b-8917-bd377116f215 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Mathematical analysis of machine learning algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.712751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.712751Z digest=sha256:1315fdca88320c625384f349f197519d9b826b91dbfed84790e7bf9538c63bca

Observation 1fbb9cad-42c1-4b29-9652-ca3470c1e0c0 · outbound

This paper cites Probability: theory and examples, volume 49.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Probability: theory and examples, volume 49

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.056031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.716671Z digest=sha256:a893c20443b1b6c03645d53c394f0ccd779497b19d90b788927c97b14b94c70e

Observation cd360758-56b3-4fd2-a23a-6f26d30c2e42 · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 34

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:45:31.043147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.720489Z digest=sha256:4271f6f178ae69cd8309a7e18ed9aa3d178dd85f63ae3e3703211b06023189c5

Observation 44af0ef1-f44d-4789-b130-ec622744a4b5 · outbound

This paper cites If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.028048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.723796Z digest=sha256:07c4d982e684ff9893948390248a648643e93a7d02c75a41622d4cd7b29e842d

Observation 5c9b9790-6d33-471b-99e3-7bde2da9912d · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.015628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.727319Z digest=sha256:5b9a0f056c68f832e7809b04c98b5b7c4305bf707a35c74d30a7eb3e189799f5

Observation 8af02ebc-195d-427a-b337-ca2f411af61f · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.004160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.731168Z digest=sha256:be3528540879d1d38dae2433bf96c946a95e8508e9640ab1b71845b9acdfb792

Observation 5c19b0e8-7570-4206-92b9-a2fd8c9ccfca · outbound

This paper cites log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh).

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:30.992308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T05:45:30.735139Z digest=sha256:767aec2048c135462be00f0487d432c09d2684f19f1fcfdc395fa43fb77574f0

Pith citing papers

Observation 39b49e20-785e-4b2e-9bf3-6d436a00e7db · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.833402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:d54bc684cef7358c27560bd8abd9bbdc1321a5699984a21e215b7da91494172e

Observation 12e287ce-b650-4ece-9a3a-d938f677967b · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:49.164337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:6cc9483afb0e4acc9e3f70426dbeb95e27a5f9a26e249ba2ad98ce3ca1156d62

Observation 979eb267-7af1-4b5b-94d9-ca1c7e9f0ca1 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:456be4df711ba816263f9f2ef1ea8b53f761bc53aad120d6405f800b7c4e618e

Observation d546229b-d332-4831-9195-e3e3a43148fb · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.651378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:38:11.785469Z digest=sha256:d9b827bf76df431ab074be72d1cc2e4473730d7df5c33cdee210b92db86db2d6

Observation ab9dbc9c-ef0d-4135-bed8-94aa1b062247 · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:31:41.386191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:31:41.386191Z digest=sha256:78a6086a0066ee2f818b1c0d7eece87bfb2d48a5a8e8475b475b9862d1bd7338

Observation fee230e5-cb3d-4ff1-923f-61f1c48c44b0 · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.425338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:3a5dd27d94781e5cfb06fbfd914f6e047b2fa0ed03aa90dce6665ff3bf799061

Observation 37134517-1422-4099-86b4-023ea435cf36 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.627224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:828b4d520271a0e007940c9219715108f40d2a6c6270bc37cad2fb3ce73e201a

Observation e1a2eaa2-1aba-4f20-8b48-d2dcdff3366a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 164

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:245d8aeafc4b1c20046b057f28e35c48332089eeafaf67d9ea0310858c59ec03