Pith. sign in

Paper Citation Record · LEDGER

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

As of 21 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 3 inbound Pith citation observations for arXiv:2412.02186.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.02186 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T23:51:14.742680Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T16:55:35.743040Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:57:30.210583Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy28
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 571d9ada-efa4-413d-aee3-db6433039085 · outbound

This paper cites Many-Shot In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Many-Shot In-Context Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.424206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.424206Z digest=sha256:f718be90473a28ac566c63e4c1f5f207bd73be9c8d859a0d3735ed07148a9a41

Observation e47c8b24-a67f-4591-9381-d5f344352b46 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.430263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.430263Z digest=sha256:be9bd52fe2c52480796ef8e45dd624eea43667ed45c6ebd89d484cdd1a22817f

Observation 33426ff5-1c28-4759-a1d7-684092f572f7 · outbound

This paper cites The Internal State of an LLM Knows When It’s Lying.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding The Internal State of an LLM Knows When It’s Lying

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.043963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.435374Z digest=sha256:13a0061910468325ed77dc49abe6f8ca60cea2fa5a3aeb83e3ad80d1f3ff4096

Observation f31be695-0c92-41b1-a431-4a1cd0c933d5 · outbound

This paper cites What makes multimodal in-context learning work? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1539–1550, 2024.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding What makes multimodal in-context learning work? In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 1539–1550, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.024541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.440438Z digest=sha256:a188beb1c0949db700726ec2e0b62e8f8594431cbb3384c3a06c68dd42b05f79

Observation 3000c2ad-479f-4229-90cd-8e93ba1caf01 · outbound

This paper cites Capera: Captioning events in aerial videos.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Capera: Captioning events in aerial videos

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:16.005161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.445729Z digest=sha256:55cd3db6cc0ce55732cb827764587480f55f0231aea8f160479739ee5aacea78

Observation 8655438b-64f3-45f2-91b0-76a5803b0250 · outbound

This paper cites In-Context Learning with Long-Context Models: An In-Depth Exploration.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding In-Context Learning with Long-Context Models: An In-Depth Exploration

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.451019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.451019Z digest=sha256:1a9ba1b785b0c2916073a67542da1cd7ed5bdca8cea19fe079747cfa2ccf9c90

Observation ee3543f8-7fc9-4aa6-9745-3bbdb8eb4e3d · outbound

This paper cites Language models are few-shot learners.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Language models are few-shot learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.456647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.456647Z digest=sha256:8a67d09560a6fc1de4fb0147525b845e191c6e88c58641410707d0634d9d5d70

Observation 82644989-1306-45ca-be67-c56ef8439b7d · outbound

This paper cites Understanding and im- proving in-context learning on vision-language models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Understanding and im- proving in-context learning on vision-language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.973461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.462278Z digest=sha256:a3bd95ae646ae931e75e6990fcc1fad44010e1ca68046c2902f018c0e85d0208

Observation f1548042-8d3a-451f-9496-311d957d6ac3 · outbound

This paper cites Mmict: Boosting multi- modal fine-tuning with in-context examples.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Mmict: Boosting multi- modal fine-tuning with in-context examples

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.954702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.466947Z digest=sha256:5770f4070774f81976b468069fd022015bc945d9da2ed2c9ada8870a0daf6f5c

Observation fa70ddb4-8f31-4b81-9b25-cc5f0a1962e0 · outbound

This paper cites Learning to retrieve iteratively for in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Learning to retrieve iteratively for in-context learning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.937127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.471843Z digest=sha256:b2994bf670b3d1569a5e7a7f3dbc627c42298f29a05fe260a82dd28a2dc02125

Observation a63fb2ee-c1c2-4ce4-a031-02c78c5757f2 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.477210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.477210Z digest=sha256:7ce8c6d6c4399d465739f053c8d5b3e15457d26ccccde908aa949b66e4c300fa

Observation 7c8af262-f323-4d8c-8962-4287cfe170a1 · outbound

This paper cites Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Don't Hallucinate, Abstain: Identifying LLM Knowledge Gaps via Multi-LLM Collaboration

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.482501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.482501Z digest=sha256:4d8f84c8c82aaa63e831c43361679c82583ccc8f4b6bad3036048bb3717fca80

Observation dab4aa2f-36ba-4c70-83a7-ca3c7d0a3ec9 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.487621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.487621Z digest=sha256:292f2e4552fbb0295bc7a5f3be625a5b5a9a916c2aa8bb65ec769bdaa111256e

Observation 2f78e198-0ccd-4c71-a409-8ae6baf53d97 · outbound

This paper cites AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding AIM: Let Any Multi-modal Large Language Models Embrace Efficient In-Context Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.493064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.493064Z digest=sha256:dc8644bed49295534a30e62ef02dc2472d3bdcbb10f251b5a998b8c0b5d48e29

Observation bcf62a95-fa19-4264-965c-386a44b6ccae · outbound

This paper cites Weinberger.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Weinberger

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.919967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.498535Z digest=sha256:50a84c0cbc909cbb15780b636edba7da68eb2a131ffbdc8f83c070eb536b215b

Observation 42a85aca-2d02-4db1-b09a-1cb6e1fa9c68 · outbound

This paper cites How well does GPT-4v(ision) adapt to distribution shifts? a preliminary investigation.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How well does GPT-4v(ision) adapt to distribution shifts? a preliminary investigation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.901497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.503179Z digest=sha256:e3af501be040a06683e94ab7c3f496bd918c0ae4aee6ca800c2407ebc6bbe297

Observation db435614-fb1b-4f4e-8870-c915208545e7 · outbound

This paper cites Pitvqa: Image-grounded text embedding llm for visual question answering in pitu- itary surgery.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Pitvqa: Image-grounded text embedding llm for visual question answering in pitu- itary surgery

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.883044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.509300Z digest=sha256:6598c7b264a2c294338f60f5368b8740b338703d5586f5c562ca45b7e4c4fa20

Observation 504544ab-74df-4598-93d0-4d31396b693a · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen- Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.514432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.514432Z digest=sha256:ac93019a2a950f971bf4424593695bc97ee75d802ffea68c8f20c03975a1f721

Observation 3970faac-4f6e-4c54-86b0-060254f2a212 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.526422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.526422Z digest=sha256:eb24191862c6d6de214cec02e65e40c7e03dbe3fba1f309e0da06dd660fe3d58

Observation 864c8ac3-f57e-4fd3-badf-2acae5754f66 · outbound

This paper cites Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Look Before You Leap: An Exploratory Study of Uncertainty Measurement for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.531598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.531598Z digest=sha256:da51554dfd0b55bd19213e5bf20debea27b53ddb6f160c9675e15e2eeea664a5

Observation 61478545-56a7-4cb3-8b02-75b7f09e2e4e · outbound

This paper cites an unresolved cited work.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T23:51:15.848646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.537478Z digest=sha256:a6d3aea683c632cad7992eaa70f0228bd9fe61237f2eb0a6bdd64edec36eb531

Observation e2aa7efa-c620-4152-9fd9-fa0620b8ca96 · outbound

This paper cites Language Models (Mostly) Know What They Know.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Language Models (Mostly) Know What They Know

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.542704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.542704Z digest=sha256:ef3786988f96492fddfcacbbce3fbfbefd2c5e9008180e22f88a6f0154f5814e

Observation 375bf7bb-c4e3-4eba-9375-9a88f36604b4 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.548402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.548402Z digest=sha256:c5e4776156752c2dccb779752cc7a4b6132025c4d4437787fc0bd61950567b74

Observation c1fdd8d8-a53b-4420-9e14-93c13df9f5f5 · outbound

This paper cites Confidence under the hood: An in- vestigation into the confidence-probability alignment in large language models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Confidence under the hood: An in- vestigation into the confidence-probability alignment in large language models

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.829784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.553667Z digest=sha256:bb788cba4f88b6f154481f93a90ecc3c322848d0e20179687e1266e06abd85be

Observation f80988a6-f330-4a3a-9922-fc1b8fdabbca · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.559194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.559194Z digest=sha256:3ad29bc7cd3566ca3d61c15acdb09c02a9112fb9986a823ef5bd107af187b0d6

Observation 6495a6b6-7781-4c06-bf5e-08b541625e84 · outbound

This paper cites MIMIC-IT: Multi-Modal In-Context Instruction Tuning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding MIMIC-IT: Multi-Modal In-Context Instruction Tuning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.564829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.564829Z digest=sha256:93dc0246bbd555f27881edd60c528fa3ea4414bfd0fdef4c86dbb9d99b7f6234

Observation 258063b0-5951-40f0-98f3-28390500df8f · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.570356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.570356Z digest=sha256:e33195902d3e7f172cdad6821bdcefc7f03978f9265da22bc88189bce25c188f

Observation 91b4ecd1-b7fb-462b-a482-d877b727e94d · outbound

This paper cites Sports-qa: A large-scale video question answering bench- mark for complex and professional sports.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Sports-qa: A large-scale video question answering bench- mark for complex and professional sports

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.575466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.575466Z digest=sha256:fb14f19b83c56f598b8505a4942ab5e5500440cc6ef43f814e2daaa493ea317b

Observation 60bb38e6-2b8d-449f-a3d6-ad47ab765ad8 · outbound

This paper cites Inference-time intervention: Elic- iting truthful answers from a language model.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Inference-time intervention: Elic- iting truthful answers from a language model

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.809890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.580480Z digest=sha256:a214cb451e423570779834078f52b1ad96dd4d2fc7c5de12152084b48dccb853

Observation 06fa6ffc-64dd-40ed-b4fa-34ec3e5a0151 · outbound

This paper cites Mvbench: A comprehensive multi- modal video understanding benchmark.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Mvbench: A comprehensive multi- modal video understanding benchmark

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.790084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.585269Z digest=sha256:280b7553fed09c3de350b87e41f1049a30e33b63f394ceccc5d0c28675dd08c4

Observation 4a89e339-044f-4ab6-ade5-500e47ce825e · outbound

This paper cites How to configure good in-context sequence for vi- sual question answering.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding How to configure good in-context sequence for vi- sual question answering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.773944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.590480Z digest=sha256:403ef4f3c53eda504ed847f45063ca0b9e4368ac83769cf1e844e7b6569ccb3a

Observation 767d76c1-ffb4-4d80-8dae-fab78d8f0b50 · outbound

This paper cites Teaching Models to Express Their Uncertainty in Words.Transactions on Machine Learning Research, 2022.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Teaching Models to Express Their Uncertainty in Words.Transactions on Machine Learning Research, 2022

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.758154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.595521Z digest=sha256:528e72f5249080d62a1e691646b604860130b6ea360969e3bf920da96a43b68e

Observation bb12d7a7-39a1-4c44-8b6a-e77ffad36758 · outbound

This paper cites se2: Se- quential example selection for in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding se2: Se- quential example selection for in-context learning

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.741506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.600669Z digest=sha256:584a94b01503aa0d3036e84a63287fd2c8f9c01aac6f4749982044efd761ab29

Observation 438170e9-4897-429d-99a3-811a3693fd8b · outbound

This paper cites Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Oryx MLLM: On-Demand Spatial-Temporal Understanding at Arbitrary Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.607317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.607317Z digest=sha256:9906907f7eb9f134aab33b981d2767bc69d5e4d50dffd7d61f139e4f30a904df

Observation a7f95e16-6273-416c-8332-f42adac9c9c8 · outbound

This paper cites Factual Con- fidence of LLMs: on Reliability and Robustness of Current Estimators.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Factual Con- fidence of LLMs: on Reliability and Robustness of Current Estimators

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.724043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.613427Z digest=sha256:1fcd1148245b04eb0221ca7030d93243ffa2d490328da6662299d73d30965d5c

Observation e1f1f03f-6b27-4f62-90f9-ac8275618929 · outbound

This paper cites Lingoqa: Visual question answering for autonomous driv- ing.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Lingoqa: Visual question answering for autonomous driv- ing

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.704720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.618912Z digest=sha256:6886300d99eac2f1e9e08008a4b8ae30e316ee1a0e880b8a4c66d45c8beba170

Observation 3575e72b-1000-41be-891d-b974fd6a9fc6 · outbound

This paper cites Drive&act: A multi-modal dataset for fine-grained driver be- havior recognition in autonomous vehicles.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Drive&act: A multi-modal dataset for fine-grained driver be- havior recognition in autonomous vehicles

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.685780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.623845Z digest=sha256:7e0b7b7df18f8119daec6e8a114228a6cf2a349d2a5a899477bae342703b2499

Observation c8ff2fc7-1db6-4411-9cbc-24528c0605cd · outbound

This paper cites Animal kingdom: A large and diverse dataset for animal behavior understanding.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Animal kingdom: A large and diverse dataset for animal behavior understanding

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.667641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.629925Z digest=sha256:af4805e4987cb77289f581e1a96d8dc619ea156087136f09c7feef70bc4b1728

Observation 087b1459-d397-449e-8f73-b5ec6697eb69 · outbound

This paper cites LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LLMs Know More Than They Show: On the Intrinsic Representation of LLM Hallucinations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.635385Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.635385Z digest=sha256:2e611a473e4bc180aa6fc419ba708e0ca442b3eadd1af9d829ca8f90b8d16108

Observation a49c6b7b-75af-4b44-b808-9ca6a5d90e8c · outbound

This paper cites Perception test: A diagnostic benchmark for multi- modal video models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Perception test: A diagnostic benchmark for multi- modal video models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.649825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.641705Z digest=sha256:be44ac18c6a79ad27e83874fec746ac3d43ad2fcd90408e8a03d2e4047355fc5

Observation 2c35c90e-15bc-4919-a061-3be3835ff96f · outbound

This paper cites In-Context Learning with Iterative Demonstration Selection.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding In-Context Learning with Iterative Demonstration Selection

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.647222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.647222Z digest=sha256:310741593f90a061978b661a108f5381495574071935f5980ab322c1d054d547

Observation e20c2581-681c-4fa7-b21f-b5a83a25ef3b · outbound

This paper cites Sentence-BERT: Sen- tence embeddings using Siamese BERT-networks.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Sentence-BERT: Sen- tence embeddings using Siamese BERT-networks

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.626230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.653373Z digest=sha256:a550474cab500cd70f47b1358e2b987188c46070bf04f0b737038cc1285e524a

Observation 9456c65a-abeb-453f-9fa3-3abb9d4f8dce · outbound

This paper cites Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Few-Shot VQA with Frozen LLMs: A Tale of Two Approaches

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-11T23:51:14.966231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.658083Z digest=sha256:e757e00ea4bc8402718fdbd55be8370c5fc390b8d637c01d7d20aa03567c15a4

Observation ef26cba9-22bb-486a-93c9-e966d44a68b3 · outbound

This paper cites Real-world anomaly detection in surveillance videos.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Real-world anomaly detection in surveillance videos

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.607434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.663210Z digest=sha256:94e5db4210d9276717acb912c8cf67d2140f3af6cdd565905f25d057af8736c1

Observation 2d6afd20-42c2-4392-a48c-3defe167753e · outbound

This paper cites Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Gemini 1.5: Unlocking multimodal under- standing across millions of tokens of context, 2024

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.584029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.668124Z digest=sha256:74ba50d13c0035f4a0b9bb719e8d8aff5c3dffc8b0c133a1f54a844f120b1ccb

Observation 10496ad8-4900-4dfe-8467-c58125151b09 · outbound

This paper cites GPT-4o System Card.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding GPT-4o System Card

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.672870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.672870Z digest=sha256:65ff995a030183159d8f05e1a2894c2d2bfcc36775d1532c2650cd9c1ee159ff

Observation bb241294-e3d6-40b9-8504-ad15d6b363e9 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.678177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.678177Z digest=sha256:02437c5b4f2b74f7add93797633b0621a19ac03687f3dd44255702a4b82e5454

Observation 95673d8d-49be-40ca-a667-b052887eba29 · outbound

This paper cites Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Bayesian Example Selection Improves In-Context Learning for Speech, Text, and Visual Modalities

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.683313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.683313Z digest=sha256:42ac8426db8ee69f2bbb2e930e8c7b49392555ccc026f2fe75babee522017a1d

Observation a15826be-eb33-4850-ad1d-cc11be042225 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.688603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.688603Z digest=sha256:756619ff1bf105e3523d8316bb833fde7cbfe7a482dc6d9ae98e60a7f3b1076f

Observation 08aec839-854c-44c3-9d1a-872afa2c601c · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Next-qa: Next phase of question-answering to explaining temporal actions

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.563368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.694586Z digest=sha256:9e44b0c35c72e81f6c392b45972f27da636598ed7e858e1f8a596e5283ba39e3

Observation a1ea448f-1d56-4f2d-8e69-5a5ae2f056a3 · outbound

This paper cites Can LLMs express their un- certainty? an empirical evaluation of confidence elicitation in LLMs.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Can LLMs express their un- certainty? an empirical evaluation of confidence elicitation in LLMs

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.542701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.699913Z digest=sha256:7af761de2e869fb713ab1c70616e25937dde683088221f4692277f1360cfb7e0

Observation f66e0c66-f136-414e-805d-a286813a179c · outbound

This paper cites From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.705316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.705316Z digest=sha256:e8c5fec663434fefa2049de7f34a1f1c1414d043336b282a29f8dad64f26fd3b

Observation 711784b5-2852-4338-8b72-966363f59a0b · outbound

This paper cites Improving the Reliability of Large Language Models by Leveraging Uncertainty-Aware In-Context Learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Improving the Reliability of Large Language Models by Leveraging Uncertainty-Aware In-Context Learning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.710709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.710709Z digest=sha256:33551f88076617bbcfb17290fa8784e7f21e72edb0878944a09122224ef36b4a

Observation 79fdc45f-e651-48e5-a508-f5815a3f6404 · outbound

This paper cites Eliciting in-context learning in vision-language models for videos through curated data dis- tributional properties.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Eliciting in-context learning in vision-language models for videos through curated data dis- tributional properties

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.518885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.715803Z digest=sha256:ead37176b15a03a32ca141c134653d3341d70ce2b66d487117ca3ba1af17acbb

Observation dd39e91c-1384-48ed-a5e7-166b2a24d26a · outbound

This paper cites Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Revisiting out-of-distribution robustness in nlp: Benchmarks, analysis, and llms evaluations

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.497000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.721435Z digest=sha256:a6dfd1c996ffe2bcfa433ba08f6ea056da9d5d1ed5aea426e0112f63fd7791c3

Observation 61f5b66a-6e46-49c1-9da7-4c93b026bf5b · outbound

This paper cites On the Out-Of-Distribution Generalization of Multimodal Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding On the Out-Of-Distribution Generalization of Multimodal Large Language Models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.726744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.726744Z digest=sha256:71a6cde2ee65d4ffa40a843d67ddae60809a5a775af96a66e4e0f65f35df12e0

Observation cf85a32f-c478-4c30-b6fe-060d703cf856 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.731849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.731849Z digest=sha256:bfb5f208ce2c5e41ea3621bde46d68e60d80ce0e8b90127410667d8794f06eaf

Observation 431dd5b0-c683-4f39-84d9-27567ab21794 · outbound

This paper cites MMICL: Empowering vision- language model with multi-modal in-context learning.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding MMICL: Empowering vision- language model with multi-modal in-context learning

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T23:51:15.473057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T23:51:14.737316Z digest=sha256:e700cfc9d16d51295133aa4357bf8fca723d970324ff35ab8d5c8a13a3b36f19

Observation a3f20680-b46d-4f7d-b93d-a17efd63ba01 · outbound

This paper cites Vector-ICL: In-context Learning with Continuous Vector Representations.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding Vector-ICL: In-context Learning with Continuous Vector Representations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.742680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.742680Z digest=sha256:ac9e646c725b304ae5c34aeb878e5246cb952075d3512f2a029057b8edfb858e

Observation 0d9cd69f-e966-445a-98e4-db5bdc4526fa · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding LoRA: Low-Rank Adaptation of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T23:51:14.520095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:51:14.520095Z digest=sha256:2223f59d20fd41d6dfa41108246ebcbcae4af0ad9148d59923536074b73d1476

Pith citing papers

Observation a7728b32-a7da-4a48-9291-08521a020d2d · inbound

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition cites this paper.

VideoNet: A Large-Scale Dataset for Domain-Specific Action Recognition VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:20:41.154852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-08T18:35:48.379198Z digest=sha256:349e5f7af57ebd6af1589b63dde1365b5da7f7a06efd08bac8b085fb654503a4

Observation 90b8d5ec-f23a-4b42-bb89-f047dff6be94 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:06:37.273682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:a7e6febbbdd2ba33a557ff6b59f6db45c3b449ce006eeb078b86bc10d05189f3

Observation e2a0e597-70e2-429d-bb1d-f593f41581e0 · inbound

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA cites this paper.

Counterfactual Reasoning for Fine-Grained Evidence Disentanglement in VideoQA VideoICL: Confidence-based Iterative In-context Learning for Out-of-Distribution Video Understanding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:57:30.212175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T16:55:35.743040Z digest=sha256:00cba66bc468b8c5c80ae0ec07fb12823673ee68dafbdd0529869fdfd4a3dfa1