Pith. sign in

Paper Citation Record · LEDGER

AdaVid: Adaptive Video-Language Pretraining

As of 23 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2504.12513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12513 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:56.191733Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe9c2dec-853d-4ac7-80c4-21ac84cb7edc · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings.

AdaVid: Adaptive Video-Language Pretraining Hiervl: Learning hierarchical video-language embeddings

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.966485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.953474Z digest=sha256:e73b532049166547b1f1ccbf5a7e48005b434d89dfd2c13672725bc6ca148496

Observation 0aef341b-cf31-4931-b3e5-5a06b043f0b8 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

AdaVid: Adaptive Video-Language Pretraining Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.951522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.958717Z digest=sha256:484f1b760bf9f92b3829a24260669f06f1bc20c77ca20bf48c8f9ce9666a718c

Observation 44a795ac-f4d3-428b-b48a-d2a600b54f28 · outbound

This paper cites Memory Consolidation Enables Long-Context Video Understanding.

AdaVid: Adaptive Video-Language Pretraining Memory Consolidation Enables Long-Context Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.963796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.963796Z digest=sha256:6d241a7685cc0867ac5f504957a8a0bc99eb9335c5d1c6daebb39b88eee195ca

Observation 52cf30be-8cf9-44a7-9954-b313918114dc · outbound

This paper cites Longformer: The Long-Document Transformer.

AdaVid: Adaptive Video-Language Pretraining Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.969231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.969231Z digest=sha256:d4c5e81b9bf098093046e23c3311ae1c1cccdd7d2ccb3500346357084cdd66e6

Observation a9c3490f-ca98-4857-86ef-acff910beb61 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

AdaVid: Adaptive Video-Language Pretraining Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.936390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.974125Z digest=sha256:a1367d8ef1927f07410a62808b77b321bbdd366c126afc0ccfe3e8a9b4151747

Observation 6e27c4bb-9e01-49fc-90e2-efccd3fd852d · outbound

This paper cites Flexivit: One model for all patch sizes.

AdaVid: Adaptive Video-Language Pretraining Flexivit: One model for all patch sizes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.921806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.978854Z digest=sha256:f79752aef27cc88586672e911f443b873dd7766ba920e5c4e5c565eaade8df8b

Observation f473e052-b15d-4cb3-b7c0-61d167d87e89 · outbound

This paper cites Once-for-All: Train One Network and Specialize it for Efficient Deployment.

AdaVid: Adaptive Video-Language Pretraining Once-for-All: Train One Network and Specialize it for Efficient Deployment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.984188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.984188Z digest=sha256:e1f329f016e93fa7ff698cfb433c0e7377d01f042a74b3023315f0dbcf0a4319

Observation 824ac077-966e-42ea-be87-e6d81d372349 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

AdaVid: Adaptive Video-Language Pretraining Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.907252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.989080Z digest=sha256:1e0c07bc2d5030752fa3d6e05b335b7c0c04bee8fb735e1faff9c8ab927b6326

Observation 096fdd9b-bd77-4d49-a29c-37af99143617 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimiza- tion space.

AdaVid: Adaptive Video-Language Pretraining Vision transformer slimming: Multi-dimension searching in continuous optimiza- tion space

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.892774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:55.993662Z digest=sha256:037159ac46a1d32d75bf1bda518e2ee097c708b58a917b8db94e955cc62a68c5

Observation 0722223b-cbd0-402f-97d8-b8f9f094dccf · outbound

This paper cites Towards the Limit of Network Quantization.

AdaVid: Adaptive Video-Language Pretraining Towards the Limit of Network Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.998203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.998203Z digest=sha256:4acfd37bf979a5718baf52449e5e16d4c2b37988ec7264a0e6c2d1dca60983e1

Observation 05a3fe94-49c0-4002-826b-fe1afda5c235 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

AdaVid: Adaptive Video-Language Pretraining BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.003372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.003372Z digest=sha256:95d4399fbea3d0c73dab4806d927e00e1ffd4fea7bafae33a89a84334c7bd4c5

Observation 440cb46e-30f8-4f78-8fa4-889d3d75a44a · outbound

This paper cites Matformer: Nested transformer for elastic inference.

AdaVid: Adaptive Video-Language Pretraining Matformer: Nested transformer for elastic inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.877891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.008261Z digest=sha256:e2eb325848afd093d46eab60dc6e0bc86e7cfff73d99f5be34cc7922f4ad3b57

Observation 77641918-6f5d-471a-85c3-986fcb0da2e0 · outbound

This paper cites Slowfast networks for video recognition.

AdaVid: Adaptive Video-Language Pretraining Slowfast networks for video recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.861335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.012750Z digest=sha256:9ab648f9f83a90d41e796b3f70a79d70971f60bd147efe515abded71db3d1046

Observation 05e8a850-a9b9-4757-bc7a-853394a9cc16 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

AdaVid: Adaptive Video-Language Pretraining Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.846268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.017739Z digest=sha256:ab0a55f67962258e35efd0b5e6c488b62a059bec6f9e32e17fca3b5c3f7ea944

Observation e3809cd4-506c-4386-91d2-8d6271a610d4 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

AdaVid: Adaptive Video-Language Pretraining Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.832277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.022435Z digest=sha256:ab9ed9720d901a1e857ce60ec4f07bbbd8127f087db841388e299fe1a5dadde8

Observation da3a14b4-915a-41bd-924d-3e905235e210 · outbound

This paper cites Dynamic convnets on tiny devices via nested sparsity.

AdaVid: Adaptive Video-Language Pretraining Dynamic convnets on tiny devices via nested sparsity

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.818036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.026815Z digest=sha256:e42e0282460824269b7c8dd904c39157f8f07a60075434ffc8706ac5102503b6

Observation bcaa72c8-997d-445c-8839-ea0a049f07a6 · outbound

This paper cites Transkimmer: Transformer Learns to Layer-wise Skim.

AdaVid: Adaptive Video-Language Pretraining Transkimmer: Transformer Learns to Layer-wise Skim

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.031204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.031204Z digest=sha256:6354a2f18500cfa9577e1049b96ab3b88af85e5900addc8025e8894b9029862b

Observation fccdba26-2254-439b-892c-2250a6d5162c · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AdaVid: Adaptive Video-Language Pretraining Distilling the Knowledge in a Neural Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.035899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.035899Z digest=sha256:e611ff3073bb9d630c3fab0e901f0d91da8252bed2700d5498a32bbdd929e1c4

Observation 4640f7b3-90cc-446e-8e9c-ce294c95de7e · outbound

This paper cites Training Compute-Optimal Large Language Models.

AdaVid: Adaptive Video-Language Pretraining Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.040813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.040813Z digest=sha256:59d3d1d87ef9cca4aa3185ea2528e8ba8ca2b7b41a4f58b894246d927c05ee6a

Observation 10756d83-cb59-4617-841a-a84c00ef0b31 · outbound

This paper cites Dynabert: Dynamic bert with adaptive width and depth.

AdaVid: Adaptive Video-Language Pretraining Dynabert: Dynamic bert with adaptive width and depth

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.804032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.045191Z digest=sha256:947d3022667d699b6cd35894230b97d27e1572370afe0ab4bc617f4f8e3692bc

Observation 7a036d8c-987c-43d3-a4b4-52b26a4e5795 · outbound

This paper cites Long movie clip classification with state-space video models.

AdaVid: Adaptive Video-Language Pretraining Long movie clip classification with state-space video models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.049562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.049562Z digest=sha256:e60b25d5f7d4c2212533e1d967858f05d6c4ede878a2f5550ec1d1285c60c47d

Observation 3fbb887f-d92d-485c-ab64-82a673efe4b0 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

AdaVid: Adaptive Video-Language Pretraining Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.054262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.054262Z digest=sha256:924b2727614ff4f68fa85112dc4006aa05a313d4c9c8530e328922754babe0a2

Observation d14186c5-df10-40b1-a9b3-05541d0585c1 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

AdaVid: Adaptive Video-Language Pretraining Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.058930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.058930Z digest=sha256:4e4b14a03c03e62858b8c61b59f6fb8a0cba07005116b55934bffa14f775eac4

Observation 0615abc7-7681-4708-a7c2-e681c661cba4 · outbound

This paper cites Matryoshka representation learning.

AdaVid: Adaptive Video-Language Pretraining Matryoshka representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.780186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.063785Z digest=sha256:7316d74b6299ee7a04a7278011b306b2a900775d59fbfee1ce61b420ce097763

Observation c354b6f8-641f-409a-b30b-0c17bbe8b3ee · outbound

This paper cites Block Pruning For Faster Transformers.

AdaVid: Adaptive Video-Language Pretraining Block Pruning For Faster Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.068793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.068793Z digest=sha256:8e128b7b870f02445ffbf675293f5a0fc5d63acace67a0f8d7fbc0e6660b792c

Observation cc3b5e32-5c90-4b00-be4f-a08a413218d0 · outbound

This paper cites Resound: Towards action recognition without representation bias.

AdaVid: Adaptive Video-Language Pretraining Resound: Towards action recognition without representation bias

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.765223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.073483Z digest=sha256:be6cbcc0547a6014f2c59acdbd19c97e379b5d857d2c623d4307b8d7139cecd7

Observation ae65ca19-72ad-4e95-bdf9-61c8ab151262 · outbound

This paper cites Egocentric Video-Language Pretraining.

AdaVid: Adaptive Video-Language Pretraining Egocentric Video-Language Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.078363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.078363Z digest=sha256:4fbb50c81a317a58ad48fa56e37572e4f2d41dcdbb9072631577ec86f97f169d

Observation ce75f7db-64a8-4ce0-9e2e-526e496bf87d · outbound

This paper cites Egocentric video-language pretraining.

AdaVid: Adaptive Video-Language Pretraining Egocentric video-language pretraining

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.749579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.083071Z digest=sha256:71f61ab93b68c16870f32e78527dd0bfdb478e367d48416d44cbc83aa949b27d

Observation c51ad084-7082-426e-b1d4-73f91270ff0b · outbound

This paper cites Rethinking the Value of Network Pruning.

AdaVid: Adaptive Video-Language Pretraining Rethinking the Value of Network Pruning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.087445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.087445Z digest=sha256:1b7ae0ccd0f2345cbd1ae80b8a5154d5e713e0afe4df101417d4f06e25faa6c7

Observation b0d0ef6f-8449-4594-a162-77d548a1b18a · outbound

This paper cites Decoupled Weight Decay Regularization.

AdaVid: Adaptive Video-Language Pretraining Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.092058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.092058Z digest=sha256:3006272b94282b027a316de3b4bf3230402ac83d53fb989366251d9065ebfc47

Observation d2053245-e3a9-4b9c-93a3-30288afa8d6e · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

AdaVid: Adaptive Video-Language Pretraining UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.096502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.096502Z digest=sha256:1095a92086031f9e90507a8ee8e0f51caeb510fd2e37c6c3e45c8424d2a5a445

Observation 87c13dce-107c-453e-8ca7-06937970a679 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

AdaVid: Adaptive Video-Language Pretraining Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.735087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.101319Z digest=sha256:ac2c74100dd3cf441f233581f70e6250aa5a591a6ff78200ab131ef499dd8547

Observation d2b4998c-b90c-492c-baa6-b8160e4604e6 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

AdaVid: Adaptive Video-Language Pretraining Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.720661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.105835Z digest=sha256:9abac85d600c5ca6ea8839bfa27ef5382bf7300138a2b34eff6cc6e38da49e83

Observation 8a0b218f-5f8f-4904-8cf0-e37b93acfa7b · outbound

This paper cites End-to-end learn- ing of visual representations from uncurated instructional videos.

AdaVid: Adaptive Video-Language Pretraining End-to-end learn- ing of visual representations from uncurated instructional videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.706236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.110318Z digest=sha256:c747bcb2e49d39778ba76c8017cc00da2195eadaff3cb7d488769ddc80170a5f

Observation cc4a0554-9109-4996-9e89-d03841bb85ca · outbound

This paper cites A sim- ple recipe for contrastively pre-training video-first encoders beyond 16 frames.

AdaVid: Adaptive Video-Language Pretraining A sim- ple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.691703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.114688Z digest=sha256:c153dcc8b1e248f3ffaed644daf385f8563cf2297a79156fe9ab56ab76084d63

Observation 721ea63a-ec2c-4c0c-93db-3822f6387b3a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AdaVid: Adaptive Video-Language Pretraining Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.119502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.119502Z digest=sha256:6251193ad287377481e11181b067830cd4dfecd8260cea867ca506a7ea84952e

Observation bc351f54-06a2-4124-a937-65c6e3239472 · outbound

This paper cites SHARCS: Efficient Transformers through Routing with Dynamic Width Sub-networks.

AdaVid: Adaptive Video-Language Pretraining SHARCS: Efficient Transformers through Routing with Dynamic Width Sub-networks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:35:56.322949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.124027Z digest=sha256:1babb65cc124f687bd2c56ad591dba50ff56d0f27851cd9aefcf6ba0412bc804

Observation a1a911d8-a3d9-4c5e-938e-8a01a045b54e · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

AdaVid: Adaptive Video-Language Pretraining DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.128650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.128650Z digest=sha256:f1e889d4db68b52eac75e8cbcc9dc4523e7f4526675875370e34841dfac81294

Observation e87bcd01-65a0-40f3-874b-df2ac13f6d83 · outbound

This paper cites Q- bert: Hessian based ultra low precision quantization of bert.

AdaVid: Adaptive Video-Language Pretraining Q- bert: Hessian based ultra low precision quantization of bert

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.668159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.133416Z digest=sha256:e0bd776d3f271d5ffc400a9dab18df91ef487d1716b0b04683a76e86aa9fbb00

Observation 63612549-91f1-42c6-84b6-e2ea945978e7 · outbound

This paper cites Training data-efficient image transformers & distillation through atten- tion.

AdaVid: Adaptive Video-Language Pretraining Training data-efficient image transformers & distillation through atten- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.653140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.138106Z digest=sha256:6ac38f55bb4c1bc4cc24e505edd38e84bf2c98d575f28b720d2e246fb007102d

Observation 2f04f4e1-1728-4516-9703-d2657e51f998 · outbound

This paper cites Attention is all you need.

AdaVid: Adaptive Video-Language Pretraining Attention is all you need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.142606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.142606Z digest=sha256:6dab8c62b45ca17aad6708ef320f27dd542c7412049cba5b947753e6d7ac5e93

Observation d3a657a0-3aed-46bf-99ea-320548a525cc · outbound

This paper cites Deformable video trans- former.

AdaVid: Adaptive Video-Language Pretraining Deformable video trans- former

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.628578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.147182Z digest=sha256:3fac91ceb32398ace652cf368f40d3e36d6715f8d9265a7c9b92866947d73470

Observation 24561286-af50-4f6d-a761-985ccb83a42c · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

AdaVid: Adaptive Video-Language Pretraining VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.151992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.151992Z digest=sha256:6c938abc79cf85df9868bff193b9224875ad8da6f0fad7c2104011828959c4b6

Observation 4c970475-6e73-41c5-a6db-0b8dcfacdd20 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

AdaVid: Adaptive Video-Language Pretraining InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.157243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.157243Z digest=sha256:592c4b488a94136d867326c5881c56362778dfeeb3f60b516367ed2f9e060482

Observation 8060ef7f-6824-4eda-b5e7-a5e9ba352308 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

AdaVid: Adaptive Video-Language Pretraining Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.162838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.162838Z digest=sha256:ca2cdb9797a4ff9bf8c8a490cddc0d18ce858976979ad1a3831189c23f350def

Observation 052cd404-e35f-47d8-b9a3-ced47168b153 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AdaVid: Adaptive Video-Language Pretraining VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.167427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.167427Z digest=sha256:bb2c638313ceb80bb1cd342f02a1dd8721027c3888627310caa052ba9886b17f

Observation c6b921f7-1817-4f60-b8c9-178b349f5c63 · outbound

This paper cites Slimmable Neural Networks.

AdaVid: Adaptive Video-Language Pretraining Slimmable Neural Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.172291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.172291Z digest=sha256:0ca9e772f3155cb66d4dcd6d5f3ad8d596df641c5efe6c881a145fa2e84ecbae

Observation c19320ce-cf02-4109-856d-7fe29833806a · outbound

This paper cites Self-chained image-language model for video localization and question answering.

AdaVid: Adaptive Video-Language Pretraining Self-chained image-language model for video localization and question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.177310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.177310Z digest=sha256:c0504d4d84ce34acd8a239e879efc89912b0279abcb622cc22aa999bbaa0f788

Observation 5739293f-3212-42ca-a1a7-c8fe5a48e90b · outbound

This paper cites Poa: Pre-training once for models of all sizes.

AdaVid: Adaptive Video-Language Pretraining Poa: Pre-training once for models of all sizes

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.595138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.181965Z digest=sha256:ff4f0836c4200b0fae4e1b3b65aec9a6e4729474aff077d394855cc2d1f3bd44

Observation d9093c44-da05-4731-ab8f-e1582d8778d9 · outbound

This paper cites Mgsampler: An explainable sampling strategy for video ac- tion recognition.

AdaVid: Adaptive Video-Language Pretraining Mgsampler: An explainable sampling strategy for video ac- tion recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.186656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.186656Z digest=sha256:fa2abdec816d82599f87a7be7e911c075b0df223762b3b7dcca3fb4d47e58228

Observation 907ccd42-b7ab-41b0-88fb-8defaf1949d5 · outbound

This paper cites FLOPs computation A.1.

AdaVid: Adaptive Video-Language Pretraining FLOPs computation A.1

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.571000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-16T12:35:56.191733Z digest=sha256:e08df62f3ccd14c248801ed5b1841f9c7cde3151accfc957d73c7b67b1429281

Pith citing papers

No inbound Pith citation observations are available.