Pith. sign in

Paper Citation Record · LEDGER

AVIS: Adaptive Test-Time Scaling for Vision-Language Models

As of 19 August 2026, this Paper Citation Record lists 86 of 86 outbound references and 3 inbound Pith citation observations for arXiv:2606.11576.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.11576 v1

Coverage vector

measured 86 of 86 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-27T10:47:41.183211Z

measured 89 of 89 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T22:09:38.842587Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T00:31:39.939203Z

Reference resolution

86 of 86 outbound references displayed

  • verified exact16
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1d10113-6a0c-4d28-a8b4-b2d7739a5272 · outbound

This paper cites Qwen2.5-VL Technical Report.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.829559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:5c63a21efc6fdc3e25a904ffff16a269e1b3101c5620c01e1224000e573e09b7

Observation bd52fa1e-df50-4ad0-8d31-20f3e5bccc3d · outbound

This paper cites LLaV A-OneVision: Easy Visual Task Transfer.Transactions on Machine Learning Research, 2025.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models LLaV A-OneVision: Easy Visual Task Transfer.Transactions on Machine Learning Research, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:f6490dfd61ee9533aa52bdf56625e67a2c93ef260595bdabd2abab513a501156

Observation ec2a6cad-d544-4e94-895d-a2fbaeb50613 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.830470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c1464e40c4dbc9fa56618d3c3b62f95c6e32061ed14f0a078600685c5b6f0ab8

Observation 27228014-78f0-4134-ac44-79659ba643a6 · outbound

This paper cites an unresolved cited work.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:a34732c49f1dcc565236cd568cb97255279a6d0d38370a9e222041bae1495d0e

Observation 425fd672-8555-450b-8ae0-4d147f4bda38 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c46aac409247b5aa24e096d566f776fbe26178308b785d2726b84471862fd1d9

Observation cae7c323-968f-49b6-9fbd-5210bdf1ed4c · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:dd42797e7513ea9eab5736d7a9dae62675b1f993c7c7c1545e5f3ff2df66844c

Observation 51a2785a-af28-4695-919e-22675b4baf02 · outbound

This paper cites Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Ziegler, Ryan Lowe, Chelsea V oss, Alec Radford, Dario Amodei, and Paul Christiano

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:88bef02f9831c5d87c03fb0e11d6211a329730ec276314f1c382ca387e243c3c

Observation 8c6c65d2-ccb7-4d8f-98b7-b40acae309ee · outbound

This paper cites Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.Transactions on Machine Learning Research, 2025.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Evaluation of Best-of-N Sampling Strategies for Language Model Alignment.Transactions on Machine Learning Research, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:e49ae9b61ebabec0b72a5a30acc23abef8da8fbc57108d5a734e6bd5f96b6388

Observation 01a4a787-68f3-4871-b3a9-54fde78796d6 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:98d477978b64eaf4d88c984883fa36d59325e18b63871703978e7533ed9bd3eb

Observation 0ee07fd5-6180-4424-8c6d-503e2ba45f88 · outbound

This paper cites An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models An Image is Worth 1/2 Tokens After Layer 2: Plug-and-Play Inference Acceleration for Large Vision-Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:1694df41729eb752fbe837668a7dc6268ee1f510f72242e18fbe742f2a3839a4

Observation 38844e25-3d3f-4444-854e-b25c4fb62241 · outbound

This paper cites LLaV A-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models LLaV A-PruMerge: Adaptive Token Reduction for Efficient Large Multimodal Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:65e847093e685b0b781a1d3c81a8bbdc3bd00dab49878fd3ed09de91af3988d1

Observation e16d7e38-1593-43ee-a354-a52ae2db81c0 · outbound

This paper cites Limits and Gains of Test-Time Scaling in Vision-Language Reasoning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Limits and Gains of Test-Time Scaling in Vision-Language Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.853645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c7138090c2d77e4b619c00ce80d1b0369e77ab2a71e010e3c99cd4c7ad92f780

Observation 087cc637-eeed-4ff3-bfbd-2f3860f7e51a · outbound

This paper cites Visual Instruction Tuning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Visual Instruction Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:efe411d451687dbdb4bb6288e5bc6d6649b9b4afd336576cc5563dfeaf9e9384

Observation d3e747b1-894c-4718-afff-6ed93a32062f · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre- training with Frozen Image Encoders and Large Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models BLIP-2: Bootstrapping Language-Image Pre- training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:03a16145036cc13340f93532612e00026663bd26c4a08bfb3bc45dcf7dbcbb34

Observation 905779dc-3e25-470e-988e-11643f7f7c4c · outbound

This paper cites LLaV A- NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models LLaV A- NeXT: Improved Reasoning, OCR, and World Knowledge, January 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:66b35701f40809dcb29905c6b7b850044ad1a2133daa4705a0da4c40d42ad081

Observation 36731078-b9d7-4c09-808d-d70e45670558 · outbound

This paper cites LLaV A-CoT: Let Vision Language Models Reason Step-by-Step.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models LLaV A-CoT: Let Vision Language Models Reason Step-by-Step

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:a0c2eb076d46cb7c023d7f846aa858d51a0911471a46f14808d49ab2b282ac0d

Observation 72064742-5105-426b-8ecd-cdc7f2488570 · outbound

This paper cites Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Mulberry: Empowering MLLM with o1-like Reasoning and Reflection via Collective Monte Carlo Tree Search

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:43aaee8e59538560a0c657a450da7b3e5d8ca2d3954e853690edb000638fcec7

Observation 9f3b7609-3c84-4d53-aa03-14f915f7ad40 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:41573b8f235f81155adad8d8f27b2cb80f75bcc9a4a16b2243776df086268487

Observation 2aed1795-662e-4db5-a588-7bd1f7bf7e9f · outbound

This paper cites Puzzle Curriculum GRPO for Vision-Centric Reasoning.arXiv preprint arXiv:2512.14944, 2025.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Puzzle Curriculum GRPO for Vision-Centric Reasoning.arXiv preprint arXiv:2512.14944, 2025

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.892179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:3a22820ef0632e09cb714bd5af4262bfe2321c7947e0ec957b884bb80382d832

Observation 6e42a8d8-6041-4a40-a37e-c7654626e1af · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.894692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:9e36628830106301d27aefeb9c67057db247a85f98b58a11d7835357f3b966f6

Observation b3eb7964-fa77-4a78-b4dc-b970b3884e84 · outbound

This paper cites R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:a639bb2d74f3db1d1e34e832737e0b7d9adc6b90f47509ba62ab8fe47a65e44d

Observation 742ba8f3-c0f9-4678-bfab-e399f5d2bb6a · outbound

This paper cites When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains.arXiv preprint arXiv:2603.01301, 2026.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models When Does RL Help Medical VLMs? Disentangling Vision, SFT, and RL Gains.arXiv preprint arXiv:2603.01301, 2026

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.837527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:02533a06d38d0db8a089c3e774cc1a2365b4ceba65d71e49dc1c2e63298f6f6d

Observation e38bf516-a527-4225-8fd5-2168b2346526 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Evaluating Large Language Models Trained on Code

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.897113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c94215b9b49fdfeca8dd1075cc7ceafa02a9889af195ae128d566b901d39a5bd

Observation ab321964-d0aa-45d7-bc3f-d5d8a82875fd · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.889259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:e1e1723182f032ecabd58071e7cc9c4445ec23fd5fd8f256dd9a4e79588f46c0

Observation 299826ec-235d-4284-9d1e-6b7e704518ff · outbound

This paper cites Majority of the Bests: Improving Best-of-N via Bootstrapping.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Majority of the Bests: Improving Best-of-N via Bootstrapping

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:b18f338c8de8c98a7e4353666024a3dcf7fb26f300302e39c7f8054b379cc1c6

Observation 4c327420-b881-45b2-a2bd-b0f7136fda44 · outbound

This paper cites Self-Refine: Iterative Refinement with Self-Feedback.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Self-Refine: Iterative Refinement with Self-Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:8049bfb4325da78bcfa1a3d46a6f976e0a23e7c245d8dd434029252e5c8770d8

Observation 6d96d1cb-d588-49b6-8129-821e39e7d313 · outbound

This paper cites Let’s Verify Step by Step.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Let’s Verify Step by Step

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:be11b9a4238f873d2ea10fb06535333d01de793d635d58eef156474dc6360e6d

Observation 61c055f9-9f92-4601-92d0-aa93e4818f0e · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:b2a30990b1de10df078a2c9c159bc99b90297a4d9a361c03efe1838acba325e4

Observation 55e7a0d6-a857-4bea-98eb-601dd7cc944a · outbound

This paper cites s1: Simple Test-Time Scaling.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models s1: Simple Test-Time Scaling

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:4e0de960bd62e20f22b2976b248a0f4a532a887b7bb639469a58eeb44445a78f

Observation 1c16a5dc-dadb-4fb5-8ea1-f9021b8ff470 · outbound

This paper cites SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models SelfBudgeter: Adaptive Token Allocation for Efficient LLM Reasoning

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.855439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:5f8156fb731e095b3021e8dfe2ccca98687318169694abcb84fcaca2b5701503

Observation 3e831cad-a8a7-4145-a1a7-e22b4c8de4be · outbound

This paper cites Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Adaptive Inference-Time Compute: LLMs Can Predict if They Can Do Better, Even Mid-Generation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.848719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ecd1124f9851d4781f062246f4ad9ef3269d5210fa9fb415bb339c05d9591373

Observation ed61dc1a-580c-4023-916f-dbf868dd6c80 · outbound

This paper cites Scaling LLM Test-Time Compute Opti- mally can be More Effective than Scaling Model Parameters.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Scaling LLM Test-Time Compute Opti- mally can be More Effective than Scaling Model Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:456eca236d01756dec213edcf98e960c5ff2a8867cc971d4a7558ff216c4d1f1

Observation c0b93ebd-32e2-45cb-a59a-9438605e49f2 · outbound

This paper cites Learning How Hard to Think: Input-Adaptive Allocation of LM Computation.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Learning How Hard to Think: Input-Adaptive Allocation of LM Computation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ca24453267f9ac21ee85f8b58c892322e3492ac2b6ed858ef974bfac18d7d48a

Observation bec2ef43-3501-4268-ade0-ea680fe68fe0 · outbound

This paper cites CyberV: Cybernetics for Test-time Scaling in Video Understanding.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models CyberV: Cybernetics for Test-time Scaling in Video Understanding

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.842398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:5f2499e0bbaf7837af665aa77fbbf72a46a6701957cddff781e2a8b387565532

Observation bb97b15e-69c1-411b-b3b6-05a219fbbe90 · outbound

This paper cites Vista: Mitigating semantic inertia in video-llms via training-free dynamic chain-of-thought routing.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Vista: Mitigating semantic inertia in video-llms via training-free dynamic chain-of-thought routing

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.844988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:e1895c2dfd5a966893d26151383da1ba2614857f211fc8fdd0ac6fb0a0f02a91

Observation 37009677-e164-4331-94ee-6427dc504079 · outbound

This paper cites Learning Adaptive Reasoning Paths for Efficient Visual Reasoning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Learning Adaptive Reasoning Paths for Efficient Visual Reasoning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.847646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:f5c05c7193873bcd8372ae0d7474bf120f1ec4f97b3135bbafebb722e315b1c2

Observation 9dbd312c-ab38-4353-abba-14e53e86babe · outbound

This paper cites ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models ARES: Multimodal Adaptive Reasoning via Difficulty-Aware Token-Level Entropy Shaping

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:d8f894d4cb5de1762971253007095129afa9474f7d26c45b5a445b1957d00b0a

Observation e91a4022-da5f-4c1c-b022-f8fb5c3f7a72 · outbound

This paper cites Efficient Test-Time Scaling for Small Vision-Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Efficient Test-Time Scaling for Small Vision-Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:60556a4511917816274629a8ec4390511d3b2eb353532ca4aa416e057577284a

Observation a1a728f2-27eb-4f0e-9a5c-adbf4f408ff1 · outbound

This paper cites Improve Vision Language Model Chain-of-thought Reasoning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Improve Vision Language Model Chain-of-thought Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:628b9f40922a9a7b1996604a575bcac632417d16321b6aad66a1472f6dce80d6

Observation c2f37f97-6c6d-47ff-860a-3f66a96f996e · outbound

This paper cites Aha moment revisited: Are vlms truly capable of self verification in inference- time scaling?arXiv preprint arXiv:2506.17417, 2025a.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Aha moment revisited: Are vlms truly capable of self verification in inference- time scaling?arXiv preprint arXiv:2506.17417, 2025a

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.850204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c1755f4442a6a7642f3e163898e800196f01a41cf97a0f1af0514771c407b410

Observation a1e3b4b9-1721-4539-9413-73f061516819 · outbound

This paper cites Similarity-Aware Token Pruning: Your VLM but Faster.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Similarity-Aware Token Pruning: Your VLM but Faster

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.886204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:725bb61cc5f41f5c1759803402aa119c41aa0f6585e574fef6ba776b44f6397b

Observation cf165c29-1646-4073-9ebc-4ad66361df70 · outbound

This paper cites PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models PyramidDrop: Accelerating Your Large Vision-Language Models via Pyramid Visual Redundancy Reduction

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:c3bea5e401a80bc560c07e0d13b1a3e7b552ec2aaabc6753efadca995d754f89

Observation 72d7ac93-7673-4491-9f3a-fc9f37fad972 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:1383eb317cc3028cd7254f724f6c4e96a675613e477cac10c23fbe9f72b4d38a

Observation abf55d21-b095-4827-a0d7-7738819772b1 · outbound

This paper cites VisionZip: Longer is Better but Not Necessary in Vision Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models VisionZip: Longer is Better but Not Necessary in Vision Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:acb19a22246986e3a9247bd5812f6830c8d92bdb18c1e345d8228dea61e06a83

Observation 06901dcc-af4c-4161-92f4-676f8d42fed4 · outbound

This paper cites PruneVid: Visual Token Pruning for Efficient Video Large Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models PruneVid: Visual Token Pruning for Efficient Video Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:cdf2e0d4c0cb8c5d29e446acee36d7c2fe0ea33f2343ac7d1ca95471d8be3d53

Observation 1c4a636f-f22a-4396-bebb-15e0992cb8b1 · outbound

This paper cites Token Merging: Your ViT But Faster.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Token Merging: Your ViT But Faster

Reference 46

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:135379aa676933f10bd05a1d0ebc9a246a7dca6ff72dbaadebaed8312ad76530

Observation d5059fa4-e6ac-45b5-b1be-7c25f572cfca · outbound

This paper cites DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models DivPrune: Diversity-based Visual Token Pruning for Large Multimodal Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:bbefc40cbaf7608791787fb16f0e0f866572c343ab6effaf1d15aa1e5b153280

Observation fd85ac8e-8961-41b5-b746-a30b97f927d8 · outbound

This paper cites Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:a10cf5b379565a835bdc38cd9609949f3dc017b8d0b106579555d501cfc86bdb

Observation 5f38d9df-4fc2-4966-b07f-4e99f125a18c · outbound

This paper cites Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Beyond Text-Visual Attention: Exploiting Visual Cues for Effective Token Pruning in VLMs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:566230e1b60a2a1e38c39f200fdd3ab69b98a52f3a327672312dcef48913a601

Observation a39d5062-6d91-4620-963a-475933721c25 · outbound

This paper cites TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models TopV: Compatible Token Pruning with Inference Time Optimization for Fast and Low-Memory Multimodal Vision Language Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ec41dc24d60f24c2d3ca292356c694afac2a0292502e022490ce9a98f7801452

Observation 63bc77e3-8231-4112-a387-473ba560c1df · outbound

This paper cites CATP: Contextu- ally Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models CATP: Contextu- ally Adaptive Token Pruning for Efficient and Enhanced Multimodal In-Context Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:0dd0deb0ac8c038302b5aea018587bf22929cd0244722478469e4fb9549e5d04

Observation 3204da28-5f07-4cd9-9c40-87eec7f451f0 · outbound

This paper cites FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models FlashAttention: Fast and Memory- Efficient Exact Attention with IO-Awareness

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:d7099e7cbaecb13d74d5ea86a955c57ffb5ffae64b37c559a0f366b29f3324a8

Observation 194aa806-0e9e-41e0-ae4e-fd6a7db41eba · outbound

This paper cites VScan: Rethinking Visual Token Reduction for Efficient Large Vision- Language Models.Transactions on Machine Learning Research, 2026.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models VScan: Rethinking Visual Token Reduction for Efficient Large Vision- Language Models.Transactions on Machine Learning Research, 2026

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:7e632343052b45087d0a2ac50cb0f2fb66ff991bacda1d69f77359a77432fd7c

Observation 0d9986cf-33b2-4062-aafa-6d9daa60909a · outbound

This paper cites KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models KeyDiff: Key Similarity-Based KV Cache Eviction for Long-Context LLM Inference in Resource-Constrained Environments

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:6284a6b0d6cfcfd548aec1f0d854c775b7c722856b0e9168b3b3fed93838ec34

Observation 31df5339-a59d-47c3-afc2-3584911f23d8 · outbound

This paper cites VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models VLMEvalKit: An Open-Source Toolkit for Evaluating Large Multi-Modality Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:106525f1ad47a04fe187539965b6ca635d90bd7a50ba30854fdeb0109dce9517

Observation 0abf0c17-b79d-42be-9b60-6ff67d19e93c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:b54aa41d8c4abef646b9ebdc1ec739866a0de4ed410a2f0ed75424729726c462

Observation 3d5a536d-7141-40b1-bc52-b4c520926909 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? InEuropean Conference on Computer Vision, 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems? InEuropean Conference on Computer Vision, 2024

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:e8e4c7aeb0a6010c18a86658ddd88835b9c6fd26e03b7798979bdd0437fc4821

Observation c586b138-e026-4719-9f58-780e3cd94b57 · outbound

This paper cites Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.Advances in Neural Information Processing Systems (NeurIPS), 37:95095–95169, 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Measuring Multimodal Mathematical Reasoning with MATH-Vision Dataset.Advances in Neural Information Processing Systems (NeurIPS), 37:95095–95169, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:b08f21d9db94258613c79ac3ae4212e3542e428bdebf1c245e8b5547661eca40

Observation 9dcd9a28-6285-40c7-ab2c-6ee507bdc2e1 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models DocVQA: A Dataset for VQA on Document Images

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:77d7eedb0186dd96ca7e9cff7b0f372c836b8c32492b64a25467634f7f50c657

Observation 3bcb69df-7035-4877-bc02-b85a5c0027c7 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:4c01319d6084a7f402d64cedf869a6c9e714936dbead7a51b50eae79588a95f6

Observation c373c94d-748c-47b3-9de4-a5b9312d4020 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ca7190ee2a08b5bd4441c72c49340be2d2053657aa69af2b9f95bd0e08d66fdc

Observation 1b58cd25-3106-42f1-bd92-3a042c2a9c8b · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models? InAdvances in Neural Information Processing Systems (NeurIPS), 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models? InAdvances in Neural Information Processing Systems (NeurIPS), 2024

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:3c4c2505fadfd8c9b73da21238443e68d235f4cfb3786be1c40bb0698220ee86

Observation 4102869c-37cb-482c-94e8-bedbbb2cf7ea · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player? InEuropean Conference on Computer Vision, 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MMBench: Is Your Multi-modal Model an All-around Player? InEuropean Conference on Computer Vision, 2024

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:dc0832c78f895ae78b28f6081cd0c03876b2e39018e6c7a6b8407b52cb24c087

Observation 93d8c24e-56df-432a-80cd-d16399b2fed4 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:5116de5cf4a3e68d3ba71de132c3bbbd5f0370aed1339783aab13b715e65b567

Observation 65262a67-54f0-4467-bfa2-bcbf099eb45e · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:99293faf9ee4af9879affa088aeef6c068ad1c3fffd1a237ab143ce4eedfa38b

Observation 8e8e5a5b-8fa7-4e66-8029-d71d64dd13a9 · outbound

This paper cites Smith, Wei-Chiu Ma, and Ranjay Krishna.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Smith, Wei-Chiu Ma, and Ranjay Krishna

Reference 66

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:93206d155d8f15f9b1f2b1bcb2b47a76515094742e4602a7bc91b74aae9caa3f

Observation 4ffeb757-ee6c-42c6-845b-eb4e9dc242ba · outbound

This paper cites Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Traceable Evidence Enhanced Visual Grounded Reasoning: Evaluation and Methodology

Reference 67

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:8ce9f7f904db74ab6f3ca51839e09d68e65a10b2e6cd7d0326123ed7f54e324c

Observation 349039f5-68e8-4699-bd8d-c0d466c4a2c7 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ad7f9cd70f3c2a5f323756fbc25230c5580a13b151fd64b395bffe945409dfd6

Observation fef393f5-4bfb-48dc-a064-20bb0c0b9bc2 · outbound

This paper cites TempCompass: Do Video LLMs Really Understand Videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models TempCompass: Do Video LLMs Really Understand Videos? InFindings of the Association for Computational Linguistics: ACL 2024, pages 8731–8772, 2024

Reference 69

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:dc3e1f5dc5846268f8ecd1a614d507be96f573ac4922a0f3521aff6a6c0888da

Observation 3b3bce9f-06b3-445b-b472-df869a9fbc17 · outbound

This paper cites Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Towards Video Thinking Test: A Holistic Benchmark for Advanced Video Reasoning and Understanding

Reference 70

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:ec6c6070d717821c9307104f41f6e8a4eac3d5d3e450fc280b2b3693cadc2d87

Observation fb99e5a9-a171-4409-9dc0-eb817966d233 · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 71

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:0d1945d123dffd7995bc49d272edba3ada7616ec1477a1009b55fd63d3ae3a9d

Observation 58316217-3316-4300-bc9b-2a9080b7e29d · outbound

This paper cites Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Q-Bench-Video: Benchmark the Video Quality Understanding of LMMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:5058f63437341ff136fdb7aee9a7b8fe222f5c883daed1b54efd4b0325bf961e

Observation 477122f4-1492-47ec-b288-ce6ac919f074 · outbound

This paper cites Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Video-MMMU: Evaluating Knowledge Acquisition from Multi-Discipline Professional Videos

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-03T08:17:45.846147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:f72b3707d3ca7acac8d8022aed57edfeeebf739b4935f180424331ec025bb9e1

Observation cdf5bef9-e05c-4e55-a5ba-be5961c06ac5 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:da98e913469178a2381753a091d29e47532f7972232003d84c1f81aa66cc6c1d

Observation f1b05748-a5e5-4f6a-b77b-ce0da3055a87 · outbound

This paper cites OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL Cycles

Reference 75

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:f09837543605952948c5e7d5a41b008274d110f06727b02387de8d4bc250305c

Observation 0781584f-13ad-4afd-8480-cde6783bf8a0 · outbound

This paper cites Perceiver: General perception with iterative attention.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Perceiver: General perception with iterative attention

Reference 76

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:51cf0ed38e6b26465693f543963efe0d171dd3c435306a82de7425af33732cf2

Observation 80c56837-1c06-4d64-87bd-666e635e9582 · outbound

This paper cites LLaV A-Video: Video Instruction Tuning With Synthetic Data.Transactions on Machine Learning Research, 2025.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models LLaV A-Video: Video Instruction Tuning With Synthetic Data.Transactions on Machine Learning Research, 2025

Reference 77

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:938cd00b88915682093d37dfcab13a743bffa658eed0ef08255ed8d3f42ee679

Observation 1c76ad6c-71e5-4935-b2f9-16bd3eaa3c37 · outbound

This paper cites A- OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models A- OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 78

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:539a15223c4e8e0b269a7ec58418449435796c66b5be44e632ac6292fe4a8cb9

Observation c8324f67-9fb9-4970-adc8-9d5e3a03ce92 · outbound

This paper cites A Diagram Is Worth A Dozen Images.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models A Diagram Is Worth A Dozen Images

Reference 79

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:4a83effd35498bbda1b364db131f94c915458e71a002d05fda62db189e91b63a

Observation 4730b2f8-e46a-4d3f-96de-10976aa182b1 · outbound

This paper cites Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Are You Smarter Than a Sixth Grader? Textbook Question Answering for Multimodal Machine Comprehension

Reference 80

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:bad9463d0390f396bf2662dce4be44d19dbab63f4db38f683dad4423ee198d94

Observation 5f1b29ca-cfde-4359-9103-f79fd6369aeb · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 81

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:f33edeb09c8baf2add8fa4d6ba834756f6f6935576946cae3757162eaec926bf

Observation 6aae95e4-c377-4639-8cdb-bc4ade857cd2 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.Advances in Neural Information Processing Systems, 35:2507–2521, 2022

Reference 82

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:b4771333fbc057ab3d33ea6bc334a4a5ee18c21eb0c3ebb63ddcf7000e7af794

Observation ba7e04dc-8421-4f85-b33f-42c64933907a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 83

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:1c2bd95df6f1aeb1b926ccd0c5fdd67d75c0f0f7325e51ea50ac5a637d61910a

Observation ed2d3af1-df27-4b17-8d27-b92274981a9d · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 84

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:adb04e03261dc58d211141bde70f16e542d2579c412ac07a2554f2db6f9c6468

Observation e0973a0c-117a-4aa3-81ee-87c8776e8ae8 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 85

Resolution
malformed identifier
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:67485027293444b25881570d867b34f340f17fcf2db8127c35a42ee207b03a96

Observation 300c80a8-6768-4b64-8bad-cb125859067d · outbound

This paper cites Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects.

AVIS: Adaptive Test-Time Scaling for Vision-Language Models Guidelines: • The answer [N/A] means that the paper does not involve crowdsourcing nor research with human subjects

Reference 86

Resolution
unresolved
no resolver link, observed 2026-06-27T10:47:41.183211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-27T10:47:41.183211Z digest=sha256:7f0d1cecac36645390d333ab924643901db82be327c20b86711e1170986237f5

Pith citing papers

Observation 94581dcf-32f0-4a7d-a91d-e740d1b659df · inbound

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling cites this paper.

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling AVIS: Adaptive Test-Time Scaling for Vision-Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T00:31:40.050482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T00:31:37.971062Z digest=sha256:5cff71b1b60905dd7131ad82ac5e3a77308c072cc46bdea14b75f95eb68c33b3

Observation 100e795c-0d12-452d-b2fe-f0ce1c4abafa · inbound

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling cites this paper.

It's the Decoding Format, Not the Perturbation: Auditing Consistency-Based Selection for Vision-Language Test-Time Scaling AVIS: Adaptive Test-Time Scaling for Vision-Language Models

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T04:18:03.076604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:18:03.076604Z digest=sha256:eba948496784455e8f3e60ea7ea5f91f6a07f5c92e36c0a0aad4f44dc8368b0b

Observation 0e7fd8d1-21cb-4921-a874-0d840380841b · inbound

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression cites this paper.

Not All Visual Tokens Are Equally Safe to Remove:Consequence-Sensitive Visual Token Compression AVIS: Adaptive Test-Time Scaling for Vision-Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T22:09:38.842587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:09:38.842587Z digest=sha256:ba2fe3501d1469d9d1f64d42a9bcad1f37461c130c80e0bfa2999a9a4ac2dabc