Pith. sign in

Paper Citation Record · LEDGER

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering

As of 9 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2502.09573.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09573 v3

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:03:13.329200Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de2ca75f-24bb-4753-b803-a9624a78f9df · outbound

This paper cites GPT-4 Technical Report.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.196965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.196965Z digest=sha256:21212178d4860e01d9b349fdd31f444aed9777b254b1432f41d7abf823f37d72

Observation 1624cbe1-9f77-4ade-bdfd-9f68932708a3 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Flamingo: a visual language model for few-shot learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.203362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.203362Z digest=sha256:54a897899b061a584506366630c219adfe5361d782740c51ef8b29730beefc56

Observation 8ddb77f4-17e0-45ba-b278-4f49f0224a65 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.209644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.209644Z digest=sha256:98ce843d9e98b75bdeac805b34c416dbde825831e0ae05ef549f2193e40a0eea

Observation b3f95510-baa5-443b-9ed1-d0c4c216c43a · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.216250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.216250Z digest=sha256:52210597e5929db5e018485a9a7a3d89d91c1c1a3ed23f386a175f3b60610213

Observation 67d1ec17-df35-401b-844b-1cc8ae526129 · outbound

This paper cites A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering A Multimodal CNN-based Tool to Censure Inappropriate Video Scenes

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-07T21:03:17.669138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.222137Z digest=sha256:d234e65fd839926b3ad63b654d6d141e7b15b8afad953eb5c2e5395906b5150f

Observation 17e20f24-31b8-4b5c-8da7-44d4284838d7 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.228207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.228207Z digest=sha256:de8eda70548656f5dd2ea9e4f3e476cd3ee32f0036db95fcca2c52b60cf42d01

Observation db718e0c-6c29-4e3f-9505-7265c13a891c · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.234882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.234882Z digest=sha256:a77511fc3046456c33cfd2c3658fb3be177fc2ca0d49eae9712f664842bde3c5

Observation 789aac87-32e9-4c96-934a-497f7fa53db5 · outbound

This paper cites Deep residual learning for image recognition.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Deep residual learning for image recognition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.241015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.241015Z digest=sha256:dac7fcd69be33667f8f7146c0fd4d4e9f85398ee66aa96839478878b8c16772f

Observation 090a0efb-7a03-4b0e-89fe-23ca011b8ccd · outbound

This paper cites H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering H., Zhang, P., Zhang, H., Yang, J., Li, C., Zhong, Y., Wang, L., Yuan, L., Zhang, L., Hwang, J.-N., et al

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.246481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.246481Z digest=sha256:7a53e00d12e96b8eee4d00a6a35afb0aedb3baf4cb1e1349a390e6a9dd338425

Observation face6fb8-7606-4124-98d4-77ae618e30fa · outbound

This paper cites an unresolved cited work.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.251781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.251781Z digest=sha256:186de691ccc43c88ce98996634decf8920378b7a7e6fc5d34de240e672b1341e

Observation a0bc587c-2b28-4d79-ad1a-0651955a31b0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.257128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.257128Z digest=sha256:c7ad1684e63eb26438758c09ded5d69df6219899131b75d6193efb0b349fee14

Observation c85518da-c75f-43c9-ab7e-5cb99c1ad88f · outbound

This paper cites Openai official api documentation, 2025.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Openai official api documentation, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.873073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.263148Z digest=sha256:3c0ca28c8a966b1f96e3195c2bfd31b528c775e5da345c1e8c076e75ad6b0037

Observation f802c8c8-efde-455f-a883-d7b3fb7cf566 · outbound

This paper cites J., Mahmud, A., Sobuj, M.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering J., Mahmud, A., Sobuj, M

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.852117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.268323Z digest=sha256:8ad865ae3dae798f54f4407d1aac985b9db37c520c693cf8e118a25988f018d7

Observation 45ade3e0-9fd1-444f-aeac-847b27b071e8 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.273301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.273301Z digest=sha256:898808fd8a03a50b67d9e49efdc8c4d057d1649d48d406028c0db72f6268d22a

Observation 8c7641b8-503c-4d55-ab5d-42edf16a08e7 · outbound

This paper cites Will the \ 1 trillion of generative ai investment pay off?, 2024.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Will the \ 1 trillion of generative ai investment pay off?, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.798891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.278477Z digest=sha256:48db1fd60a7d952246d6bdc33f9f79b756ad91356b856e5679c493bbedfe158a

Observation 631c7893-09fa-4324-a0d3-12c3c8e99927 · outbound

This paper cites Video understanding with large language models: A survey.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Video understanding with large language models: A survey

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.283982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.283982Z digest=sha256:c00f88e8a9a152e029ce0e186911448b8344eb022ca161979bef35bf128fa7ed

Observation c6fdfcdf-b3b3-4dc0-9a83-71a1ed37e48a · outbound

This paper cites PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering PEER: Expertizing Domain-Specific Tasks with a Multi-Agent Framework and Tuning Methods

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.289380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.289380Z digest=sha256:73fd276ed828493c180b08056e674f84d156375eac751f6641cbd64cb1dae225

Observation be6885a3-3740-4ae4-bd0e-13f3421c0db6 · outbound

This paper cites Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Truly Multi-modal YouTube-8M Video Classification with Video, Audio, and Text

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-07T21:03:13.409766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.295499Z digest=sha256:a10cdd31b497a0bb52fac71c8952b8b1c6da970e8839ff20a37e72a71702fe0c

Observation ff8f4093-4c2c-4c11-8553-c46bc9b15223 · outbound

This paper cites and Nawaz, T.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering and Nawaz, T

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:03:17.780330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T21:03:13.300703Z digest=sha256:9c4dcf7990a0e8c508073c165c051983c16e9e5b0a701cfeabcfcee7f2f29fed

Observation 43e1c9d8-a0fa-4e48-878b-a258b5620dc3 · outbound

This paper cites Fine-tuning Large Language Models for Domain-specific Machine Translation.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Fine-tuning Large Language Models for Domain-specific Machine Translation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.306086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.306086Z digest=sha256:248b1f2c8fc723d0a12f8110f6f2cdb022ac74519b64b83bee3660d7b5f41de3

Observation 64a935d5-5239-436a-aaaa-507bab5120fd · outbound

This paper cites write newline.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering write newline

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.311663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.311663Z digest=sha256:7fc4b3fd0d89ee4a5239137457f63d6517dcb3fa7984b7f073c3b79e0aaaa028

Observation 56d84102-f72a-4a7e-b68f-f0a97fef3345 · outbound

This paper cites @esa (Ref.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering @esa (Ref

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.318053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.318053Z digest=sha256:ee2a701b81fc9a92ab6eac8abc41bfa062cde0ab5eb4fea67ecb78fedfdd4a37

Observation 955c383f-d17d-4541-a83a-83e4be449b2a · outbound

This paper cites an unresolved cited work.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.323649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.323649Z digest=sha256:54fe44eb18b28b45bc7fa1a9185c751526eaecbbb71651aa33c424bc73173e0c

Observation 2f342397-317a-445a-b690-75ab64be3803 · outbound

This paper cites DriveLM: Driving with Graph Visual Question Answering.

Optimizing GPT for Video Understanding: Zero-Shot Performance and Prompt Engineering DriveLM: Driving with Graph Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T21:03:13.329200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:03:13.329200Z digest=sha256:39b49bc39adecb9e26ce70362644779338dc918efccd3c1b21019a144975b8d5

Pith citing papers

No inbound Pith citation observations are available.