Pith. sign in

Paper Citation Record · LEDGER

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

As of 13 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2412.02186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02186 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:51:14.742680Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:55:35.743040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:57:30.210583Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 571d9ada-efa4-413d-aee3-db6433039085 · outbound

This paper cites Many-Shot In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Many-Shot In-Context Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.424206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.424206Z digest=sha256:2f40c1e4ee8594be2e1e727af001588f8d9293c4e8d119897c9af880cbec1302

Observation e47c8b24-a67f-4591-9381-d5f344352b46 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.430263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.430263Z digest=sha256:c5242f433dc4489bdb505fce336b1337e3919e2ec4b116b29022cee4de5ce793

Observation 33426ff5-1c28-4759-a1d7-684092f572f7 · outbound

This paper cites The Internal State of an LLM Knows When It’s Lying.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding The Internal State of an LLM Knows When It’s Lying

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.043963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.435374Z digest=sha256:838b24fbc41a9c8d4ced86d0982c0db99efc76a287f98ac185da197fc6bf2633

Observation f31be695-0c92-41b1-a431-4a1cd0c933d5 · outbound

This paper cites What makes multimodal in-context learning work? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1539–1550, 2024.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding What makes multimodal in-context learning work? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1539–1550, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.024541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.440438Z digest=sha256:33b7224155fea345e153ad85bf20b0ea634cc00aadf5cf87e847dfb76ca6a2b5

Observation 3000c2ad-479f-4229-90cd-8e93ba1caf01 · outbound

This paper cites Capera: Captioning events in aerial videos.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Capera: Captioning events in aerial videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.005161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.445729Z digest=sha256:2df43157e0c7a35ef10906e7fe980611ccfcc70711b087bf6920a451cccf8a8c

Observation 8655438b-64f3-45f2-91b0-76a5803b0250 · outbound

This paper cites In-Context Learning with Long-Context Models: An In-Depth Exploration.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding In-Context Learning with Long-Context Models: An In-Depth Exploration

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.451019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.451019Z digest=sha256:253fb7427b3536b5de3576752e1714c61350b5c7b510ac5d8276cebca2b5e06e

Observation ee3543f8-7fc9-4aa6-9745-3bbdb8eb4e3d · outbound

This paper cites Language models are few-shot learners.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.456647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.456647Z digest=sha256:d7a70b590acecc34ba5b7c372292fb01e9c79327d95b14d018940f7559726d8e

Observation 82644989-1306-45ca-be67-c56ef8439b7d · outbound

This paper cites Understanding and im- proving in-context learning on vision-language models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Understanding and im- proving in-context learning on vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.973461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.462278Z digest=sha256:2107caa5516ad1c07baf5a8901b23b905a4b463d41061c0ca6c1b419e7e7e2f3

Observation f1548042-8d3a-451f-9496-311d957d6ac3 · outbound

This paper cites Mmict: Boosting multi- modal fine-tuning with in-context examples.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Mmict: Boosting multi- modal fine-tuning with in-context examples

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.954702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.466947Z digest=sha256:f407e6d1a356e0a3c67490a7e53a1b47485f5479cefd6cd9d19daf895e6193e3

Observation fa70ddb4-8f31-4b81-9b25-cc5f0a1962e0 · outbound

This paper cites Learning to retrieve iteratively for in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Learning to retrieve iteratively for in-context learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.937127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.471843Z digest=sha256:0c9d156c86af8613636a3c30960eaad38a7a219632bb3eb02998599bd7e1f643

Observation a63fb2ee-c1c2-4ce4-a031-02c78c5757f2 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.477210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.477210Z digest=sha256:d3c8a7738428cdf3cce880e2a314bc63f975649053bedbfbbe758a2b42fc0601

Observation 7c8af262-f323-4d8c-8962-4287cfe170a1 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.482501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.482501Z digest=sha256:0065c9cab4fcc2048edffb08dba8e3435c8ee923cce7ff1bc740039e325c7a2e

Observation dab4aa2f-36ba-4c70-83a7-ca3c7d0a3ec9 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.487621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.487621Z digest=sha256:ae254ccb9c9f688e8d16e4a1a93f95bce7aee703efba33bea515e1466653c84b

Observation 2f78e198-0ccd-4c71-a409-8ae6baf53d97 · outbound

This paper cites AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.493064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.493064Z digest=sha256:b5fe396c86b3d70980d4d827881f67818752bc5203aee5b6e49a6b8b79e6e287

Observation bcf62a95-fa19-4264-965c-386a44b6ccae · outbound

This paper cites Weinberger.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Weinberger

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.919967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.498535Z digest=sha256:c4629f7bfd456b99ca7242f88a8f9c588c0859e907375796027887bc4d9bf3c0

Observation 42a85aca-2d02-4db1-b09a-1cb6e1fa9c68 · outbound

This paper cites How well does GPT-4v(ision) adapt to distribution shifts? a preliminary investigation.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How well does GPT-4v(ision) adapt to distribution shifts? a preliminary investigation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.901497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.503179Z digest=sha256:575968ce4df9bbc928d2f62a61f2491965b5955a984a5ab6740a1cd3cb1c1c49

Observation db435614-fb1b-4f4e-8870-c915208545e7 · outbound

This paper cites Pitvqa: Image-grounded text embedding llm for visual question answering in pitu- itary surgery.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Pitvqa: Image-grounded text embedding llm for visual question answering in pitu- itary surgery

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.883044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.509300Z digest=sha256:791e666a05c335eb8c1827e36589ebc9558b23511e33f61f0eba44728f23fa14

Observation 504544ab-74df-4598-93d0-4d31396b693a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.514432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.514432Z digest=sha256:6ca63cefabe3a7f7bedfd3ead722b7a2740a48c62a77dbab60ef90bc75af2dd6

Observation 3970faac-4f6e-4c54-86b0-060254f2a212 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.526422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.526422Z digest=sha256:e657ad30fc4bf64e35826b47bbc2781c42d20a3aadd4f78f21357d402d34268c

Observation 864c8ac3-f57e-4fd3-badf-2acae5754f66 · outbound

This paper cites Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.531598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.531598Z digest=sha256:0d758bdd357976ecb8c7e148bca34b25df6952191e7255b9d6040bb0e17b9f3c

Observation 61478545-56a7-4cb3-8b02-75b7f09e2e4e · outbound

This paper cites an unresolved cited work.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:51:15.848646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.537478Z digest=sha256:c8c131174ef6f05b2bc37b55da6fd741b007f8e66d5cad557d5960f9b616ac15

Observation e2aa7efa-c620-4152-9fd9-fa0620b8ca96 · outbound

This paper cites Language Models (Mostly) Know What They Know.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Language Models (Mostly) Know What They Know

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.542704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.542704Z digest=sha256:b25016a4118b2814d936d9abb877eec23214518201b80bc8316584c8e6d9e7c3

Observation 375bf7bb-c4e3-4eba-9375-9a88f36604b4 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.548402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.548402Z digest=sha256:dbd47932d43a65d47481da610ae32050b76ba3349128b85de9cd10de83019232

Observation c1fdd8d8-a53b-4420-9e14-93c13df9f5f5 · outbound

This paper cites Confidence under the hood: An in- vestigation into the confidence-probability alignment in large language models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Confidence under the hood: An in- vestigation into the confidence-probability alignment in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.829784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.553667Z digest=sha256:fa163ff3f5adfa4ce65fa63c73d29a4c1478313d14accda3d5a8901491d6b80e

Observation f80988a6-f330-4a3a-9922-fc1b8fdabbca · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.559194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.559194Z digest=sha256:03216c737240c6aa6f163432836089669c49cd94962c9836345c144681ac48f6

Observation 6495a6b6-7781-4c06-bf5e-08b541625e84 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.564829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.564829Z digest=sha256:2e6b9ca79ae906faa8dd7e4574f4471e25e9510fa2fc2b208811e711727f2d79

Observation 258063b0-5951-40f0-98f3-28390500df8f · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.570356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.570356Z digest=sha256:b133cb35cc453524b36a2d4b9acffba825e1ebad7247a6788824a37cecd8d3f5

Observation 91b4ecd1-b7fb-462b-a482-d877b727e94d · outbound

This paper cites Sports-qa: A large-scale video question answering bench- mark for complex and professional sports.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Sports-qa: A large-scale video question answering bench- mark for complex and professional sports

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.575466Z digest=sha256:95261ce2ba7331ddcf54121dc6ec3ff7b6147ba7bf4bd4b7f226b5135ef97971

Observation 60bb38e6-2b8d-449f-a3d6-ad47ab765ad8 · outbound

This paper cites Inference-time intervention: Elic- iting truthful answers from a language model.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Inference-time intervention: Elic- iting truthful answers from a language model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.809890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.580480Z digest=sha256:414ea4e436e2665c4f7410487c4a7bab265f2554d0aac37f26b1c58180bfe156

Observation 06fa6ffc-64dd-40ed-b4fa-34ec3e5a0151 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.790084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.585269Z digest=sha256:8f876fea1d2ab97918c3e79f7de8bb0d68bc8b1badebd313f5322a4ca893ba2f

Observation 4a89e339-044f-4ab6-ade5-500e47ce825e · outbound

This paper cites How to configure good in-context sequence for vi- sual question answering.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How to configure good in-context sequence for vi- sual question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.773944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.590480Z digest=sha256:d5d943f7cfe4e5f8b46cff5d9b4ddf8ddb1d2eff0734365639d1d11436cb15c0

Observation 767d76c1-ffb4-4d80-8dae-fab78d8f0b50 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.Transactions on Machine Learning Research, 2022.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Teaching Models to Express Their Uncertainty in Words.Transactions on Machine Learning Research, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.758154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.595521Z digest=sha256:1abf1ee75f1b6094abd82e7cd4cdaac466689ee743f893c98e82d7c5f0365d43

Observation bb12d7a7-39a1-4c44-8b6a-e77ffad36758 · outbound

This paper cites se2: Se- quential example selection for in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding se2: Se- quential example selection for in-context learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.741506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.600669Z digest=sha256:d79643d916c6cee8acc0076dccb6ce144276db50e440b472bc0c0bed55b339b1

Observation 438170e9-4897-429d-99a3-811a3693fd8b · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.607317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.607317Z digest=sha256:ac4993f44e70abc5cbf8475fafaf4250a33002ac8b724e794f2fcf3559090863

Observation a7f95e16-6273-416c-8332-f42adac9c9c8 · outbound

This paper cites Factual Con- fidence of LLMs: on Reliability and Robustness of Current Estimators.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Factual Con- fidence of LLMs: on Reliability and Robustness of Current Estimators

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.724043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.613427Z digest=sha256:669c36758de216b0db038ab6e8a9fc2a801ac83d942cb982052e9e0b4e0ebbe1

Observation e1f1f03f-6b27-4f62-90f9-ac8275618929 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driv- ing.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Lingoqa: Visual question answering for autonomous driv- ing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.704720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.618912Z digest=sha256:b01fdbf74b6405277df87ef0e8ffb8f29818b4fe4baa97491e31f38456738a1d

Observation 3575e72b-1000-41be-891d-b974fd6a9fc6 · outbound

This paper cites Drive&act: A multi-modal dataset for fine-grained driver be- havior recognition in autonomous vehicles.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Drive&act: A multi-modal dataset for fine-grained driver be- havior recognition in autonomous vehicles

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.685780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.623845Z digest=sha256:826b6700eb3124bada0b21dca72f52a4b3ce60b6653f92921d630fd5cf0194f0

Observation c8ff2fc7-1db6-4411-9cbc-24528c0605cd · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.667641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.629925Z digest=sha256:98a11b3860c01583a457af8244f405e5637af43d35eae4cf7ed3b19fc23b8534

Observation 087b1459-d397-449e-8f73-b5ec6697eb69 · outbound

This paper cites LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.635385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.635385Z digest=sha256:e841a60cda23158921cd94f36c610bc95e10bf2fb59d5838aaa86215c994eaca

Observation a49c6b7b-75af-4b44-b808-9ca6a5d90e8c · outbound

This paper cites Perception test: A diagnostic benchmark for multi- modal video models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Perception test: A diagnostic benchmark for multi- modal video models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.649825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.641705Z digest=sha256:c41b3fb4f886673e7c8105a3228cb6bfb859efde624b426e49d864d4924c58cd

Observation 2c35c90e-15bc-4919-a061-3be3835ff96f · outbound

This paper cites In-Context Learning with Iterative Demonstration Selection.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding In-Context Learning with Iterative Demonstration Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.647222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.647222Z digest=sha256:99fc32f6b6813360e9158ba71b90aa3ca6de9000c87142cc2865e53a340fa197

Observation e20c2581-681c-4fa7-b21f-b5a83a25ef3b · outbound

This paper cites Sentence-BERT: Sen- tence embeddings using Siamese BERT-networks.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Sentence-BERT: Sen- tence embeddings using Siamese BERT-networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.626230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.653373Z digest=sha256:ce44d60a26e4b3db0a878f89a683c3a6b41ac4c4cccc87c5bb3706e70019908c

Observation 9456c65a-abeb-453f-9fa3-3abb9d4f8dce · outbound

This paper cites Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:51:14.966231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.658083Z digest=sha256:cb9d21e40a06ae078d167ebe470129edaaaa3f981a0cc6a1221e074f4281cf05

Observation ef26cba9-22bb-486a-93c9-e966d44a68b3 · outbound

This paper cites Real-world anomaly detection in surveillance videos.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Real-world anomaly detection in surveillance videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.607434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.663210Z digest=sha256:57e18d3607b804e4c7d5da1050073938a8fb04a4e33849b0760cc91745f5abb2

Observation 2d6afd20-42c2-4392-a48c-3defe167753e · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.584029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.668124Z digest=sha256:834c64819b2154200660035aa5815317642d9a01e9c2ff87d51a0246c70e0e51

Observation 10496ad8-4900-4dfe-8467-c58125151b09 · outbound

This paper cites GPT-4o System Card.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding GPT-4o System Card

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.672870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.672870Z digest=sha256:17ea8a66b990d6bac5d80b03d6a2d0dc292fb68c6a904f90a784960e3a450ccc

Observation bb241294-e3d6-40b9-8504-ad15d6b363e9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.678177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.678177Z digest=sha256:264585eb78801cb7770fb36b8a1c3b404483a68e477e776f245e66467133def1

Observation 95673d8d-49be-40ca-a667-b052887eba29 · outbound

This paper cites Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.683313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.683313Z digest=sha256:3d80b94c9b39d04d835554115042fcf976ac39809332d1c7e337e36c2f6c65e0

Observation a15826be-eb33-4850-ad1d-cc11be042225 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.688603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.688603Z digest=sha256:a533ff34d09f7d368505a846a283af355d9856e1e1a63ad6c4986ebad36971cc

Observation 08aec839-854c-44c3-9d1a-872afa2c601c · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.563368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.694586Z digest=sha256:63f1ee0cbe0126c0c86fcf7611e60f5bdb8c37aa35edd2959c4fedee477ff334

Observation a1ea448f-1d56-4f2d-8e69-5a5ae2f056a3 · outbound

This paper cites Can LLMs express their un- certainty? an empirical evaluation of confidence elicitation in LLMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Can LLMs express their un- certainty? an empirical evaluation of confidence elicitation in LLMs

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.542701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.699913Z digest=sha256:398e3d3cddd98a3e34a401c7dcd21f9bc8fad2e89bf0becfafda14e9bd911291

Observation f66e0c66-f136-414e-805d-a286813a179c · outbound

This paper cites From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.705316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.705316Z digest=sha256:48d86a551a24f54682169c131a0936e70cd6a17f15e2fa1e1de56f557bc717f2

Observation 711784b5-2852-4338-8b72-966363f59a0b · outbound

This paper cites Improving the Reliability of Large Language Models by Leveraging Uncertainty-Aware In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Improving the Reliability of Large Language Models by Leveraging Uncertainty-Aware In-Context Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.710709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.710709Z digest=sha256:ba8085c7334b1a10d61303b0d681309a3f9a8dbbae2ebb65b4e7ac33c8548e0b

Observation 79fdc45f-e651-48e5-a508-f5815a3f6404 · outbound

This paper cites Eliciting in-context learning in vision-language models for videos through curated data dis- tributional properties.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Eliciting in-context learning in vision-language models for videos through curated data dis- tributional properties

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.518885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.715803Z digest=sha256:cbd934ce9dae1dc4fc83d0ff50681caaa9d4d905360e670a95e938ea26a1cd8d

Observation dd39e91c-1384-48ed-a5e7-166b2a24d26a · outbound

This paper cites Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.497000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.721435Z digest=sha256:6d9918bfe61ef181ab5538ca78af7cbc1a911dca84d2df9261063c4dbcfd6d11

Observation 61f5b66a-6e46-49c1-9da7-4c93b026bf5b · outbound

This paper cites On the Out-Of-Distribution Generalization of Multimodal Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.726744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.726744Z digest=sha256:dc24af89541e359e993ca03f1c8e0d68d92faf3640766f38d6d3bfbc9e077d40

Observation cf85a32f-c478-4c30-b6fe-060d703cf856 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.731849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.731849Z digest=sha256:d05449ec35657648a9485c77acea4c6d36dc2b9f4ccf54840428bf2fdc459f71

Observation 431dd5b0-c683-4f39-84d9-27567ab21794 · outbound

This paper cites MMICL: Empowering vision- language model with multi-modal in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding MMICL: Empowering vision- language model with multi-modal in-context learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.473057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T23:51:14.737316Z digest=sha256:d5f8add12e287737aec87c00e888d076f9a11093fc4246ace87bb50e26218547

Observation a3f20680-b46d-4f7d-b93d-a17efd63ba01 · outbound

This paper cites Vector-ICL: In-context Learning with Continuous Vector Representations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Vector-ICL: In-context Learning with Continuous Vector Representations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.742680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.742680Z digest=sha256:fbe994f9e4f561923e7ce1898c848d615ae41e66310a651a1c74d879431124ac

Observation 0d9cd69f-e966-445a-98e4-db5bdc4526fa · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.520095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.520095Z digest=sha256:ae5bd7f37232bbd2faf4801b55abaf579ab52b066cc18ca6d44f2139b2cc528e

Pith citing papers

Observation a7728b32-a7da-4a48-9291-08521a020d2d · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.154852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:a68749ebe9b3cd4d571c2e65cdec22d2ea4557b042824870d2ff730350449dbe

Observation 90b8d5ec-f23a-4b42-bb89-f047dff6be94 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.273682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:7b2f5d11427e17579ec849482e20137e9339d21c61826caa184888e80e7fd3ea

Observation e2a0e597-70e2-429d-bb1d-f593f41581e0 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.212175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:36e6c21091af712c23516f9f9f262e378987e5e4987d967be41f7253de06a404