Pith. sign in

Paper Citation Record · LEDGER

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

As of 8 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 8 inbound Pith citation observations for arXiv:2506.07235.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07235 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:45:30.735139Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T05:31:41.386191Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T13:43:28.625652Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy10
  • unresolved27
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 26949dd0-946c-4b20-8aae-73a130af3245 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Blink: Multimodal large language models can see but not perceive

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.201843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.587797Z digest=sha256:19f5cacd2c046fe212a597f98eeb1f3c99035cf1be439c69352ab76e6eda1586

Observation 252017fe-5670-4cb6-b23d-e17e2c40e8b6 · outbound

This paper cites V∗: Guided visual search as a core mechanism in multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V∗: Guided visual search as a core mechanism in multimodal llms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.188988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.591770Z digest=sha256:8cc2a1894abb1e1b6e6be20af7dce50f15ec420982c2cf95186ded70c229e5e8

Observation e674edb9-6e54-4a7f-aa85-33d81f32411a · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.595949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.595949Z digest=sha256:d917d34f09cd6082034e627d948ac153628084401aa0ebdaff1210bf29deee0f

Observation e5b51c26-c286-46df-9016-544557ae8b24 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.600959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.600959Z digest=sha256:a3b0afe9619258b862d33734e2911fc269e3664c45b84a8c6061283b32666ab6

Observation bbb19b7a-cbac-4017-9c6e-430cd949d879 · outbound

This paper cites Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual CoT: Advancing Multi-Modal Language Models with a Comprehensive Dataset and Benchmark for Chain-of-Thought Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.605338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.605338Z digest=sha256:9fcb3c144677c670f398b4e56818a95f63ba557da0e5223dddd9c6076e2220cd

Observation 52e12acf-5b2a-4255-8ca5-94afd5e420d3 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.610059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.610059Z digest=sha256:2ba9cf2f1dd91637fe1125232ae48336fae50b6fa20f7813c9cacaece4003d1a

Observation 5d8a4253-37a2-4eec-abcf-5c91400e7576 · outbound

This paper cites Visual program distillation: Distilling tools and programmatic reasoning into vision-language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual program distillation: Distilling tools and programmatic reasoning into vision-language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.168232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.615181Z digest=sha256:4c227ae25f4252a845e449c0d27763a472e5f0b2b5b6fabf24f2705bcbc5cbef

Observation cbcf874c-d05b-42ec-a1f6-e7368d67ad0b · outbound

This paper cites Visual programming: Compositional visual reasoning without training.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Visual programming: Compositional visual reasoning without training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.619311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.619311Z digest=sha256:646c57afd439c4546f857da81f8d84b0a4f474290f09f8ae449f469992471801

Observation 9b9772bb-683c-4563-8847-1e8b248f6754 · outbound

This paper cites Vipergpt: Visual inference via python execution for reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Vipergpt: Visual inference via python execution for reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.146028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.623036Z digest=sha256:d290619b188bbafd8887ae22d729acbd05f8dfdc87aca0b05eaac4a97e5329f9

Observation 92032de7-d6d7-48c5-9a3f-a5cf9317df0b · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.627374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.627374Z digest=sha256:c89efaba4be22a752e309476d2234ee69d0e7b03fd59291fbae497c4396990e5

Observation b523a58f-2eef-46ee-b336-3950e403d8e8 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.631557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.631557Z digest=sha256:07df02b9d74c45f47c1ca4db7ed5c493b080e682f0b01447ce1f5b221ea9504a

Observation bf3df506-9ebc-4103-a789-956e840cebd0 · outbound

This paper cites Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Scaffolding Coordinates to Promote Vision-Language Coordination in Large Multi-Modal Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.635922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.635922Z digest=sha256:5ab4823cfc694138aee26c3141c3f1f02063c785accad2b2d6922bc949f8657d

Observation dcf6d6da-1844-4d90-a1f5-fa4898a2e6f5 · outbound

This paper cites MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MM-REACT: Prompting ChatGPT for Multimodal Reasoning and Action

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.640388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.640388Z digest=sha256:9b64edaf17141c3cde26f373e2e59678af633536ad693be2fdeb16d51fb30669

Observation fc168fb3-3060-4879-b51d-a94c1a8661d8 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Making language models better reasoners with step-aware verifier

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.644103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.644103Z digest=sha256:d0b93178acfffaec96b31cc044151b761b16891745615fc1501f6abf99aeea8c

Observation cb1f1667-6ae4-482e-a7ad-06dfc7a1262a · outbound

This paper cites Let’s verify step by step.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let’s verify step by step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.648369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.648369Z digest=sha256:11f53cca8516d032d9f90030f5c8d832688e22d0db62453893e370966b869a3c

Observation 3e4fe20d-b57c-481a-b5cd-c493fd0e6d42 · outbound

This paper cites Solving math word problems via cooperative reasoning induced language models.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Solving math word problems via cooperative reasoning induced language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.109578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.652126Z digest=sha256:e093d47498baa3551f767e37614cdde923a1e9f5c8e80b6dfdb1a646475a6555

Observation fb622090-c761-4f60-8374-09194f456762 · outbound

This paper cites Let's reward step by step: Step-Level reward model as the Navigators for Reasoning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Let's reward step by step: Step-Level reward model as the Navigators for Reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.655336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.655336Z digest=sha256:1ac7f3537cd89084b7a0d48bc10fc3941bbf236f82e3e9f05714731be8d6d1cb

Observation 2e143f65-c3c8-497e-892e-f6efc4614f71 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.658831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.658831Z digest=sha256:e6b995a8413d5b0ba5a31093b120cb6cf8e5eb5111ad3c420c475a016230910f

Observation 68ac5975-e95e-4064-b93e-b95e536655c7 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Fine-Tuning Language Models from Human Preferences

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.662617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.662617Z digest=sha256:72be28fdd8bbdaaa1f4c06f46d8cbc6b2e5118957556b03270f2afdee8699737

Observation 184ebd8e-b7e4-47db-80de-756321b7cf7b · outbound

This paper cites Rlcd: Reinforcement learning from contrastive distillation for lm alignment.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Rlcd: Reinforcement learning from contrastive distillation for lm alignment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.097704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.666869Z digest=sha256:68b606eeb0115e5e0b82c8d7c0a6cb67b1f1157c978aa784560445b202afb8a1

Observation 3ba739d5-16d2-4232-bc94-8f028c22e319 · outbound

This paper cites Pretraining language models with human preferences.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Pretraining language models with human preferences

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.084696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.671380Z digest=sha256:54f13681041462368dfb7ca3f4bf8f7d5129301d90df9d58112115a70b38d65d

Observation cdee5b0e-b45c-4eef-b8b3-e02f34ede260 · outbound

This paper cites Improving Large Language Model Fine-tuning for Solving Math Problems.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Improving Large Language Model Fine-tuning for Solving Math Problems

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.675404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.675404Z digest=sha256:6b51d3f8b9ac4399a62d0ddcfda97cba4c911e21e00da13d2aca4160fecb008c

Observation b221ca03-9306-44e2-b182-2b87ba767138 · outbound

This paper cites V-STaR: Training Verifiers for Self-Taught Reasoners.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification V-STaR: Training Verifiers for Self-Taught Reasoners

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.679605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.679605Z digest=sha256:522eccfd1108fe582159ff05ca13e8f70f1cbed4d316b2af6d8844581940fbd0

Observation 0e2ad5c0-c823-4ec9-b278-869e749f1dc5 · outbound

This paper cites Building Math Agents with Multi-Turn Iterative Preference Learning.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Building Math Agents with Multi-Turn Iterative Preference Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.683409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.683409Z digest=sha256:363da192103d2db80f1cf7dee4528c2ef882e1d8eb615fa80855c32ead94110b

Observation 4bde0720-d8c2-4b59-923e-cfb12b1e46bc · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification LLaVA-OneVision: Easy Visual Task Transfer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.687744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.687744Z digest=sha256:3b74f6f3b3a0266a8e974e665180e3a1ec438339fb85a6731375748cb2aef9d1

Observation 7a1641a5-f9cd-41d4-b1f7-03d615ea423b · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.691615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.691615Z digest=sha256:0e912f05735e23f8768f5e70bfbfa0a3d08d25c8e60f320f7d72a69267c9c55d

Observation 5afb27e4-9bc5-454f-a275-f37192943a9f · outbound

This paper cites Qwen2.5-VL Technical Report.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Qwen2.5-VL Technical Report

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.694953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.694953Z digest=sha256:b28d7bd782a1e7a434083f36a57584b357455a5d94edb87816510ac9fe8d6a0e

Observation f957249a-e9a8-470d-837c-08b495156d34 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.698822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.698822Z digest=sha256:a6519fdde2bf80794192a4a9322ad7f388d37fcd8cc21e8d7129d3e56d58a86f

Observation 459c0db6-b89b-4b85-948e-7185cb7681f2 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.702326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.702326Z digest=sha256:abcf499553b15faaa2119fb6ec2a8a30f0ec8846c4a4d5314a5b06673be87fcf

Observation 65f1e13d-11f2-4578-9337-f6ffec9cf394 · outbound

This paper cites MMFactory: A Universal Solution Search Engine for Vision-Language Tasks.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification MMFactory: A Universal Solution Search Engine for Vision-Language Tasks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.706182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.706182Z digest=sha256:3ae2c44ecb1f1b8bf137c9fe9b7ec946ade4a589b4dd5973d0b5e83c16abdfd3

Observation 0d54322b-c595-4edf-bbea-a1c08399fd84 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Direct preference optimization: Your language model is secretly a reward model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.709569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.709569Z digest=sha256:fb6c06ce945232b700788541317cb8a04e1abde8cbc57923a3735b581cbf5515

Observation 2c26b258-5e7b-4a1b-8917-bd377116f215 · outbound

This paper cites Mathematical analysis of machine learning algorithms.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Mathematical analysis of machine learning algorithms

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:30.712751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:30.712751Z digest=sha256:cf6c80f01de9445c7448aa5ea3aec9ca219833b01909a5459b02f1a0e9a9f9f8

Observation 1fbb9cad-42c1-4b29-9652-ca3470c1e0c0 · outbound

This paper cites Probability: theory and examples, volume 49.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Probability: theory and examples, volume 49

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.056031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.716671Z digest=sha256:32952258516940767c622465899ccc1e2796560343891150472e24697eb96940

Observation cd360758-56b3-4fd2-a23a-6f26d30c2e42 · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 34

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:45:31.043147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.720489Z digest=sha256:c7a77d575386a043ac7573575b0bb5bf0d99886d375f55e1d0c208f65575c6db

Observation 44af0ef1-f44d-4789-b130-ec622744a4b5 · outbound

This paper cites If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification If in the last definition,= is replaced by ≤ or ≥, then Xn is said to be a supermartingale or submartingale, respectively

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:31.028048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.723796Z digest=sha256:55fa93219b15e570943b4ea465feeb7531ad5051f2c03c7e807aa774e5fbb5b4

Observation 5c9b9790-6d33-471b-99e3-7bde2da9912d · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.015628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.727319Z digest=sha256:bfba14c383561abc2439c214c621fe964532e662709444285efdad6610fd4fb9

Observation 8af02ebc-195d-427a-b337-ca2f411af61f · outbound

This paper cites an unresolved cited work.

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:45:31.004160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.731168Z digest=sha256:3fec97cd4a1222f4b6e3fffd2d98c0fc4433e568daa8773bb647a924da71e1ef

Observation 5c19b0e8-7570-4206-92b9-a2fd8c9ccfca · outbound

This paper cites log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh).

Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification log V ˆϕSDPO (th | sh) Vϕ0 (th | sh) +log V ˆϕSDPO (ah | th) Vϕ0 (ah | th) # . Observe that condition (ii) implies that Eth∼R ˆθSFT(·|sh)

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:45:30.992308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:45:30.735139Z digest=sha256:4ee466dd71c9ffdbeecafd537c2cdd2e015143a1739ceffcbddab56e1c2df158

Pith citing papers

Observation 39b49e20-785e-4b2e-9bf3-6d436a00e7db · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.833402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:3a56e26dbd59dd58748c5a50ad989d445b5d76dc62821779ee594ccafffd4bd9

Observation 12e287ce-b650-4ece-9a3a-d938f677967b · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:45:49.164337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:36:42.100191Z digest=sha256:697e9440cd4ff2d92606a54a54fc975068807eba4156577f087e09b576e7d509

Observation 979eb267-7af1-4b5b-94d9-ca1c7e9f0ca1 · inbound

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models cites this paper.

OpenWorldLib: A Unified Codebase and Definition of Advanced World Models Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T09:42:23.808691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:42:23.808691Z digest=sha256:c97c60777aeb3eede4b0406423f9e60d4c193c34f5d4aa151d532d73cf44ebdd

Observation d546229b-d332-4831-9195-e3e3a43148fb · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:26:01.651378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:38:11.785469Z digest=sha256:cb2c0ea18b64403f4f640f8fa7aecf12d7764bea84fd617e915fca458cd22596

Observation ab9dbc9c-ef0d-4135-bed8-94aa1b062247 · inbound

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images cites this paper.

Test-time Scaling over Perception: Resolving the Grounding Paradox in Thinking with Images Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T05:31:41.386191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:31:41.386191Z digest=sha256:14938697c41adfa21848c30cc1ac46f772ebf8280fe87f40dd10d162754839c6

Observation fee230e5-cb3d-4ff1-923f-61f1c48c44b0 · inbound

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles cites this paper.

Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-22T07:51:15.425338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T07:51:13.362986Z digest=sha256:c19b414c47df460f2f98350ebe9e44f57f9ea2b9af8b887d77e987e0857d7220

Observation 37134517-1422-4099-86b4-023ea435cf36 · inbound

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning cites this paper.

ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:43:28.627224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:41:44.049230Z digest=sha256:67b24ab89581111f9377e5eaeaebc09da92b9c6df4cdf3c25687c4d639e0b6b4

Observation e1a2eaa2-1aba-4f20-8b48-d2dcdff3366a · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception Multi-Step Visual Reasoning with Visual Tokens Scaling and Verification

Reference 164

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:606e8caf813bf1c899313c9db1b59601fbe1c52b909289cd6188bfa706bc4d99