Pith. sign in

Paper Citation Record · LEDGER

Conformal Predictions for Human Action Recognition with Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 1 inbound Pith citation observation for arXiv:2502.06631.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06631 v2

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.636325Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:57:12.479732Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T14:57:12.816902Z

Reference resolution

32 of 32 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved6
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · outbound

This paper cites Conformal Predictions for Human Action Recognition with Vision-Language Models.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:ed921ebce03ff491b746431479ce55b784bf853a2bac5378e085c5f649d7fd27

Observation 14ddd180-af64-41a0-8915-b6c19d18bcd8 · outbound

This paper cites Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions Providing reliable confidence estimates for predictions made by deep learning models is essential in many applications

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.269711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.486097Z digest=sha256:746bb92906ebd4bfe64626bdeedfc24955b5f00c0d7e4de0a0c8a2358877ae45

Observation 937c3d2d-80a0-4e38-ba04-34cc503621e2 · outbound

This paper cites Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20].

Conformal Predictions for Human Action Recognition with Vision-Language Models Experimental settings We conduct our experiments using three video clip datasets: HMDB51 (51 classes) [23], UCF101 (101 classes) [24] and Kinetics400 (400 classes) [20]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.253420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.491485Z digest=sha256:8f7fb8ec4fa69e692ff07e6a703fb8c6837798adfcf4d088e20652a8f5596de3

Observation a4dc4bf8-b50e-4c10-974f-14b4c4860548 · outbound

This paper cites Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs.

Conformal Predictions for Human Action Recognition with Vision-Language Models Red dots mark 1/τ∗, our estimate for minimizing tail size, while green dots indicate the true optimum 1/τopt when it differs

Reference 4

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T14:57:13.222306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.502192Z digest=sha256:ddc3484948fccfd955e138e35dc29147a72ec4801d74998b418959d2e8de175f

Observation d271ec38-87e0-4102-b171-7f46843b75e1 · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.205996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.507668Z digest=sha256:7a7e4e9d222041be830857b4916ceb1cdf5667123dd1cc4dc4ce7c2fa89120e5

Observation 26d32a1e-c5df-4c3a-9808-2337a5e9288e · outbound

This paper cites Behavior recognition via sparse spatio-temporal features,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Behavior recognition via sparse spatio-temporal features,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.108366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.538451Z digest=sha256:35a56fe8d4496c4c03076601de3576b524eb72c937e307f0971aeea9f2db664a

Observation e53a7350-a19a-46af-bf66-01c5b696be03 · outbound

This paper cites Fast user-guided video object segmen- tation by interaction-and-propagation networks,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fast user-guided video object segmen- tation by interaction-and-propagation networks,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.190000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.512849Z digest=sha256:a66871ef882b47f8c2efaf3ee090d7ac021d68b8eed45f01bbb45b77eb50a7d3

Observation d623cd3e-30ec-4323-b602-6198080154fa · outbound

This paper cites Human-in-the-loop vehicle reid,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Human-in-the-loop vehicle reid,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.173741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.517839Z digest=sha256:2fabbb8d815c8fb3f1286fa4ec87b90fb862c68bfb776cabf3a1aebe5b3ae4ff

Observation db002e33-79c5-412c-be50-1a43bcc1e030 · outbound

This paper cites Surveillance video querying with a human-in-the-loop,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Surveillance video querying with a human-in-the-loop,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.157968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.523018Z digest=sha256:8c2ac465de4494253c87ad2ba61c1508f8c440bc8ee48da291c2eef94dafc092

Observation 2bf64a99-9bb0-400a-8ffc-ef1fb4a94a89 · outbound

This paper cites De- signing decision support systems using counterfactual prediction sets,.

Conformal Predictions for Human Action Recognition with Vision-Language Models De- signing decision support systems using counterfactual prediction sets,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.140776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.528008Z digest=sha256:5b4cb2a3b71cf1f7adf3568832165b4d758b183a410682dfb18a68e7d44294ee

Observation 983ee99c-06c2-4e4e-a30d-dc7291c60a6e · outbound

This paper cites Conformal prediction sets improve human de- cision making,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal prediction sets improve human de- cision making,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.124735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.532984Z digest=sha256:b494bf06584682f4d9fe0c2b35bef1fb6b2e2a76654a003ce2229e33fde5b67e

Observation b0c089b6-f903-40ce-836e-ad89ca43de0b · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Learning transferable visual models from natural lan- guage supervision,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.027326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.568095Z digest=sha256:4d537f7f57db82d23721beb96b5d64756fae9221b58c42367aa8e01af3cfb26d

Observation e0520cae-d70e-437d-9a80-e678ddd82761 · outbound

This paper cites Temporal segment networks: Towards good practices for deep ac- tion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Temporal segment networks: Towards good practices for deep ac- tion recognition,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.092733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.543548Z digest=sha256:bda4f9853eaccf7bb3a44f2b7fff95b75b30634ae6d3af96abaaec75dd2b4a3b

Observation 40a9001d-cfa5-414b-892c-87b68362b7ef · outbound

This paper cites Slowfast networks for video recog- nition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Slowfast networks for video recog- nition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.076522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.548509Z digest=sha256:5e673fa439fe10dc874cede68f8c8c707ab4853b9ac1d9ca629ae1724701662c

Observation f5c9e6ea-c327-487b-b214-588895d3b2a0 · outbound

This paper cites Stm: Spatiotemporal and motion en- coding for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Stm: Spatiotemporal and motion en- coding for action recognition,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.060036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.553338Z digest=sha256:72230d3d2db4f6ddc858c3aa6dd488a98551cb90778c39cefac54b0d3655d9a5

Observation 1c500f2c-4430-4e4f-8aa2-aea9dcb298ac · outbound

This paper cites Vivit: A video vision transformer,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Vivit: A video vision transformer,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.043979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.558219Z digest=sha256:c58cfc9874732fe4ff71341ade4212aee07bfa14e18144b1f3ca40cf6ad2b600

Observation db75da33-556d-40dd-8625-aa56000599e1 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Conformal Predictions for Human Action Recognition with Vision-Language Models ActionCLIP: A New Paradigm for Video Action Recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.562855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.562855Z digest=sha256:5e7c29d70dbdec4fae7be54508e8fa8d1ca7ae7f0f524ce00131b99043f38fc8

Observation 7f5472a1-8231-4ad2-ba2e-fc17ed0bcccf · outbound

This paper cites An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,.

Conformal Predictions for Human Action Recognition with Vision-Language Models An electronic nose-based assistive diagnostic prototype for lung cancer detection with conformal prediction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.944526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.597946Z digest=sha256:5e1c9985a983a581d22b5c001111dca0b31af1c0f7dc4e4e4e4075024f64bed3

Observation b923805b-bfda-4892-b4ee-f3ecc20ef08e · outbound

This paper cites Are foundation models for computer vision good conformal predictors?,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Are foundation models for computer vision good conformal predictors?,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.572903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.572903Z digest=sha256:dc316b4a5d359fcf312014a2ab1540e0e0bd6d65a9ba69bb8a5c90a129483208

Observation 1ed4e0fd-0643-484b-98cf-27143aa82821 · outbound

This paper cites On the rate of gain of information,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On the rate of gain of information,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:13.010707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.577660Z digest=sha256:0523b99c553f70266ab2df961dcb343355df0e7d7d1f100146bf1044eaf5ecf7

Observation a24f7d33-6532-45e0-b602-fb6ae52e3661 · outbound

This paper cites Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Selection from alphabetic and numeric menu trees using a touch screen: breadth, depth, and width,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.994010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.583346Z digest=sha256:e8bc92d411f30131c00765ff4b0da78ea38cfc10ee6f47ea9513892e62ea482e

Observation da4e27eb-11d8-4d45-bc02-038e32f85a63 · outbound

This paper cites 29, Springer, 2005.

Conformal Predictions for Human Action Recognition with Vision-Language Models 29, Springer, 2005

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.977618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.588442Z digest=sha256:4b581e23bf37be65109fa0a68c75f3123cfa7f62a81c19803ff96e9daed1ce8c

Observation 721a98cc-f580-44f3-ad99-72371ade35c7 · outbound

This paper cites Least ambiguous set-valued classifiers with bounded error levels,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Least ambiguous set-valued classifiers with bounded error levels,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.961553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.593351Z digest=sha256:5b5ed466773a2688534b5cef030a5d03dccdfec926893b94d1f4a2a66879d96a

Observation 6646e6a3-2c74-42a6-8ec3-0cdf63832d99 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Conformal Predictions for Human Action Recognition with Vision-Language Models UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.626354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.626354Z digest=sha256:547d16ce8309bd1ae707d5d748b7be4ea47c4592200f39c8244b33be4e4caae3

Observation e8f59025-c1f8-483f-9807-cf3da5364aad · outbound

This paper cites an unresolved cited work.

Conformal Predictions for Human Action Recognition with Vision-Language Models Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:57:13.237883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.496797Z digest=sha256:d7a4546355df4681ae47221b35d994c1ce02d1f4715f67efb2666a5ea27213cc

Observation b09704fd-e233-4b1d-a05d-12750ded83d7 · outbound

This paper cites Dense trajectories and motion bound- ary descriptors for action recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Dense trajectories and motion bound- ary descriptors for action recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.928675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.602719Z digest=sha256:7a7ca2d73bba619beba4a5a2d75ce5cbf94e72ff44d9fa2b597f7c274f492009

Observation bf0d9c6c-9e8d-4b62-9611-04c033e18da2 · outbound

This paper cites Quo vadis, ac- tion recognition? a new model and the kinetics dataset,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Quo vadis, ac- tion recognition? a new model and the kinetics dataset,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.910729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.607464Z digest=sha256:2ff6c7151a3b44a96af6f40ba91931043558f2f2a6aa09b34012328471ec262c

Observation 0eea96b6-72fb-453c-84f4-41be3bb39538 · outbound

This paper cites Expanding language-image pre- trained models for general video recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Expanding language-image pre- trained models for general video recognition,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.894085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.612212Z digest=sha256:c46e8b053f63f17d9753cec6bca016d2a15c9d93b36822c077a7c52c57ff14e6

Observation da2774e8-bb9f-449c-a0f4-21fb6629e4d9 · outbound

This paper cites Fine-tuned clip models are efficient video learners,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Fine-tuned clip models are efficient video learners,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.876493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.616854Z digest=sha256:9b534b6bdb82a2d2fa21fb84d799014410fb09e9bb2264d35d5e84ddebbeeac2

Observation 68c96306-e637-4d2b-8e1e-a2fb307bdab1 · outbound

This paper cites Hmdb: A large video database for human motion recognition,.

Conformal Predictions for Human Action Recognition with Vision-Language Models Hmdb: A large video database for human motion recognition,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.858847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.621764Z digest=sha256:dd8076c02fea519364b1f0c77463fb38b0e078e9a852404cba8a12007e3af5c9

Observation 3f56ef6a-da66-46bb-ae9a-b5593bf1c613 · outbound

This paper cites On sequence learning models: Open-loop control not strictly guided by hick’s law,.

Conformal Predictions for Human Action Recognition with Vision-Language Models On sequence learning models: Open-loop control not strictly guided by hick’s law,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:57:12.842209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.631485Z digest=sha256:aa244c51d8c28cdb157ed9d7092322b3647df0bb458a56d2ccfbd53342fb2ebb

Observation 6bf53bbe-ea34-459b-95a4-8b2c145abedf · outbound

This paper cites EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters.

Conformal Predictions for Human Action Recognition with Vision-Language Models EVA-CLIP-18B: Scaling CLIP to 18 Billion Parameters

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:57:12.636325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:57:12.636325Z digest=sha256:cb75e79087078dbf9bb86f03dc1e1d275f31a7f7d4e6321eff48479e2f0f1aa2

Pith citing papers

Observation 0696ad96-3d03-4398-baa9-416714377ac5 · inbound

Conformal Predictions for Human Action Recognition with Vision-Language Models cites this paper.

Conformal Predictions for Human Action Recognition with Vision-Language Models Conformal Predictions for Human Action Recognition with Vision-Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:57:12.824486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T14:57:12.479732Z digest=sha256:ed921ebce03ff491b746431479ce55b784bf853a2bac5378e085c5f649d7fd27