Pith. sign in

Paper Citation Record · LEDGER

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

As of 5 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 4 inbound Pith citation observations for arXiv:2510.00705.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.00705 v3

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:27:59.419144Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T20:30:57.931134Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-06-29T11:53:23.711158Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fd28181a-4b7a-425d-80e1-0194aefd3762 · outbound

This paper cites (2017) were sampled at 3 FPS, while videos from ActivityNet Captions Krishna et al.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs (2017) were sampled at 3 FPS, while videos from ActivityNet Captions Krishna et al

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:59.419144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:59.419144Z digest=sha256:458a737b748c4f5af5d3475110b446d45d2c522e80627fb602ac4e2b0b203a94

Observation 32918eff-f67b-4b0f-b8f4-dff201ae4846 · outbound

This paper cites Qwen2.5-VL Technical Report.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:55.997413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:55.997413Z digest=sha256:35360d1844b0ae5ac78d6fb5c13c5277265521955df91584d27006839bef9880

Observation 1977dc0e-43d8-426e-852f-1019f5773a81 · outbound

This paper cites Language Models (Mostly) Know What They Know.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Language Models (Mostly) Know What They Know

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.827613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.827613Z digest=sha256:cac477a06d3dbe1af4bf8e276533d9e4cb2a22553fda603375b0d753a7fb901e

Observation 65eacf5c-4fdb-4d56-b248-02c3a025184e · outbound

This paper cites TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs TAG: A Simple Yet Effective Temporal-Aware Approach for Zero-Shot Video Temporal Grounding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.956200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.956200Z digest=sha256:936ee8f33b79e87ba998598b1a182c9e40c99ec76e77d8f5d15d91af95aa9ec6

Observation bcfa624a-27df-4efc-9fae-659e00ae8aab · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.131368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.131368Z digest=sha256:f88d92e1aad432db25b8fedb7e2366b23b7e3fdc53831fb722feef49a9294687

Observation 968f33d1-246c-4337-bb3f-7db71a08302a · outbound

This paper cites TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs TextCoT: Zoom In for Enhanced Multimodal Text-Rich Image Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.265975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.265975Z digest=sha256:ff83bee2903371570ed2d5e2dc126a9b88728be38d18c19ae61df5634c1b8dc1

Observation b2f74308-0b8f-4cda-a923-fb5fea0b3e70 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.393319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.393319Z digest=sha256:99d2d32dfc12a04f037a65bc38f4602942fad3fdbfa44097c1851f032cc50913

Observation 229c6e97-d91c-4bcd-a2ad-e567be55af63 · outbound

This paper cites The information bottleneck method.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs The information bottleneck method

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.524001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.524001Z digest=sha256:11c65b8258946475f19e94badb09b4bde331df2b6dd2af2d3e13fed42f1f31b7

Observation edc14937-629f-47b1-b776-1df37a0fe805 · outbound

This paper cites VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs VaLiD: Mitigating the Hallucination of Large Vision Language Models by Visual Layer Fusion Contrastive Decoding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.747048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.747048Z digest=sha256:ca4843a0f9fb62996cea311f8633a881adf6c56552aaeead59d65ed317f44fb1

Observation 83e97f27-3ede-4322-9404-30dabd8d60ff · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.880798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.880798Z digest=sha256:6dc548bf930d469354367bb3c0836955640262c36aa6da9e01ef933d9e2c751e

Observation 88ca922f-510e-4504-9f08-5e73b4c37d34 · outbound

This paper cites A Survey on Video Temporal Grounding with Multimodal Large Language Model.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs A Survey on Video Temporal Grounding with Multimodal Large Language Model

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.123306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.123306Z digest=sha256:0a3f869162410d6f86b99b7189ccb522ab7a9d1a019344772fea2cccebf023a6

Observation c768bec0-c8c8-4a76-8389-60235b0a8898 · outbound

This paper cites Generate, but verify: Reducing hallucination in vision-language models with retrospective resam- pling.arXiv preprint arXiv:2504.13169, 2025b.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Generate, but verify: Reducing hallucination in vision-language models with retrospective resam- pling.arXiv preprint arXiv:2504.13169, 2025b

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.267926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.267926Z digest=sha256:46c16d88ee8b8364af8d2e15e7a32f9e0d0ba8e6fa821c335f0c36a51ba7813b

Observation 63e2b82f-acc5-4c00-8289-ba68f89b3760 · outbound

This paper cites LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs LMMs-Eval: Reality Check on the Evaluation of Large Multimodal Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.447481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.447481Z digest=sha256:898cd495bdf6d90c950ffe6ff3c15771f8fbcbe16e4cb6e1460b6b74bc29399b

Observation 4438ba32-5b43-4ea1-948c-2d2e64bf8975 · outbound

This paper cites DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs DeepEyes: Incentivizing "Thinking with Images" via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.615053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.615053Z digest=sha256:164a71a6c6bcdfbd903423b3e52040e2aa72c61acbeb0b7319ae53497000e318

Observation e29c6fd0-4f3e-4ee1-848f-d899feec41dc · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.792970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.792970Z digest=sha256:7fde323d478018a976a7f677ec496cddb6fa89cb7f671bbe933c3f5fb403e035

Observation 77e04a20-0b5d-49c9-a544-ab6deaa77d44 · outbound

This paper cites Entropy, a foundational concept from information theory (Shannon, 1948), provides a formal measure of the uncertainty inherent in a probability distribution.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Entropy, a foundational concept from information theory (Shannon, 1948), provides a formal measure of the uncertainty inherent in a probability distribution

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.913595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.913595Z digest=sha256:76e2e4ae8fc47f80935c13eef5974762194ce8f10cc75dce0275d4d19a1c1939

Observation 824337d7-b091-4191-89f2-fd29d73d9074 · outbound

This paper cites Pretrained on massive text corpora, LLMs learn to generate reliable probability distributions over a predefined vocabulary.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Pretrained on massive text corpora, LLMs learn to generate reliable probability distributions over a predefined vocabulary

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:59.081176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:59.081176Z digest=sha256:09c328f1933e75eb617b680c0fa2e5eed3e2d9a55d4e92debe7520cd814c228e

Observation 787f4816-e869-4bc5-9ded-b5ff21e63174 · outbound

This paper cites A” and “B.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs A” and “B

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:59.308256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:59.308256Z digest=sha256:5ec1eca4eec4d4ffbc3138ed59c6d86f13b8e97f9bc1fadca90161f85b61f95e

Observation e2eab76a-570b-4bac-a351-c243020a3693 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2000

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:57.642563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:57.642563Z digest=sha256:b9babdf83ffb5cdd17d92ea1d1cea4ea3b51ef6f492f43aa59749ab0dde0785f

Observation 64e6e1a3-0bf5-4063-9b05-b25003ebc6f3 · outbound

This paper cites This finding aligns with principles from curriculum learning, where task difficulty can be measured by model uncertainty (Bengio et al., 2009; Kumar et al., 2010).

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs This finding aligns with principles from curriculum learning, where task difficulty can be measured by model uncertainty (Bengio et al., 2009; Kumar et al., 2010)

Reference 2008

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:59.171856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:59.171856Z digest=sha256:2069980255d2c3c9844af5e78e34b941314cf8fc2361049fec8c741d2ef2f0c6

Observation 09861003-066e-4122-a817-4db240f692bc · outbound

This paper cites The Internal State of an LLM Knows When It's Lying.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs The Internal State of an LLM Knows When It's Lying

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:55.678696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:55.678696Z digest=sha256:b6c93dfc24962f0ea6e7e4480055afa447d59e236fa639bbf17112c712ab2613

Observation 76dbb3a0-84e8-434c-9c6f-48b3029920af · outbound

This paper cites Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Threading Keyframe with Narratives: MLLMs as Strong Long Video Comprehenders

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.359929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.359929Z digest=sha256:97d4eb16a4ad387b28ea1e211001d788f473913743fc61b05de3cac1e4b04693

Observation 1f2463c6-462a-448c-822a-86e27085b6c5 · outbound

This paper cites TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs TimeMarker: A Versatile Video-LLM for Long and Short Video Understanding with Superior Temporal Localization Ability

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.192678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.192678Z digest=sha256:22f6220da99302f290222ab9937cb6e7252ba6c61d3cdd406b765df2bedf5f28

Observation af50635d-a1d9-4977-9180-3439cbf037d9 · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:58.011777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:58.011777Z digest=sha256:e7e6ccfa68db368d556fcaa69079d8e27a1a1e6b778a399fe1ca85736aed5244

Observation 0e247db2-211e-49d4-8e58-e7bf5f68018e · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:55.803283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:55.803283Z digest=sha256:77aef85994492a028aec5fadff07ceaac0864fbc45f34c2a98ed9f40dd27c790

Observation b25ee4a3-3132-4bde-b546-d701ac7253df · outbound

This paper cites FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.481199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.481199Z digest=sha256:d343879c243f3e9711c64e6fb16392edab914c9f682b8eaa3bf12ba48951cf52

Observation a08925e4-ee69-435f-8e78-a70652459d75 · outbound

This paper cites GPT-4o System Card.

Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs GPT-4o System Card

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T13:27:56.641925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:27:56.641925Z digest=sha256:e08a7a1675c24a98a08635e51a996c4a998b137bd6d4daea6a0e29750c18c37f

Pith citing papers

Observation 658fb90c-48e9-46b5-9ec5-7ac6969bc808 · inbound

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models cites this paper.

LookWise: Knowing When and Where to Look for Fine-Grained Visual Reasoning in Multimodal Large Language Models Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T20:30:57.931134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:30:57.931134Z digest=sha256:6a65301cf8f3b3b890724a4598de085de6a7d4ffbb3e894133a37442a0a2c22a

Observation 98a8be6d-bab6-4ddb-a9ee-16a8dd9656c2 · inbound

Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines cites this paper.

Zoom Consistency: A Free Confidence Signal in Multi-Step Visual Grounding Pipelines Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:52.151779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:01:21.985632Z digest=sha256:5a7bb3aca5f5fe385172ab9c925aaab10b0f28e4ffb32f787ccbae38ede55a50

Observation d0ec393c-dffe-4757-af64-54b9d339b4d5 · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-29T11:53:23.712612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T11:51:01.769882Z digest=sha256:5bc37cb0f2d5e1ca10e9338aee1ba96f9338f47fe21c81fc91741f7b82e72432

Observation 3267df31-8efa-4278-a340-0a8425ae41b9 · inbound

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes cites this paper.

DenoiseRL: Bootstrapping Reasoning Models to Recover from Noisy Prefixes Training-free Uncertainty Guidance for Complex Visual Tasks with MLLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T12:58:10.911725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:58:10.911725Z digest=sha256:8043f788b163b347fcd66106b30cee84c880f66a0a8f30720b43d23afa99bb5f