Pith. sign in

Paper Citation Record · LEDGER

AdaVid: Adaptive Video-Language Pretraining

As of 18 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 0 inbound Pith citation observations for arXiv:2504.12513.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.12513 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:35:56.191733Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy24
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fe9c2dec-853d-4ac7-80c4-21ac84cb7edc · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings.

AdaVid: Adaptive Video-Language Pretraining Hiervl: Learning hierarchical video-language embeddings

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.966485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.953474Z digest=sha256:cf4ebd04c6b614a2fada3cf091be58f15a9a97e4149d420f6749347421e9afbd

Observation 0aef341b-cf31-4931-b3e5-5a06b043f0b8 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to- end retrieval.

AdaVid: Adaptive Video-Language Pretraining Frozen in time: A joint video and image encoder for end-to- end retrieval

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.951522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.958717Z digest=sha256:1ae6198e490b2959fa498ffb919f8f134158aaf64784c715530eb885b099df13

Observation 44a795ac-f4d3-428b-b48a-d2a600b54f28 · outbound

This paper cites Memory Consolidation Enables Long-Context Video Understanding.

AdaVid: Adaptive Video-Language Pretraining Memory Consolidation Enables Long-Context Video Understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.963796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.963796Z digest=sha256:01fe8bad62408a874967c48d0291a75caceacb7de1d37a2ac3c442e30da398dd

Observation 52cf30be-8cf9-44a7-9954-b313918114dc · outbound

This paper cites Longformer: The Long-Document Transformer.

AdaVid: Adaptive Video-Language Pretraining Longformer: The Long-Document Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.969231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.969231Z digest=sha256:f5c046481be5dd3a0a40badbc978e9605fa1a5fee3af665c2d43ddb14a0221ed

Observation a9c3490f-ca98-4857-86ef-acff910beb61 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

AdaVid: Adaptive Video-Language Pretraining Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.936390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.974125Z digest=sha256:f0099d999b99bf2130074a14da788b69609596c10e51dcc784011309f570fadd

Observation 6e27c4bb-9e01-49fc-90e2-efccd3fd852d · outbound

This paper cites Flexivit: One model for all patch sizes.

AdaVid: Adaptive Video-Language Pretraining Flexivit: One model for all patch sizes

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.921806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.978854Z digest=sha256:3d162a7d07942f403b5f1cf184e46d57dfee9627404c465098994d377ccd566d

Observation f473e052-b15d-4cb3-b7c0-61d167d87e89 · outbound

This paper cites Once-for-All: Train One Network and Specialize it for Efficient Deployment.

AdaVid: Adaptive Video-Language Pretraining Once-for-All: Train One Network and Specialize it for Efficient Deployment

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.984188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.984188Z digest=sha256:fc49940db0c51804833c597d48d6a888db02528cca441e05792009cfb475f460

Observation 824ac077-966e-42ea-be87-e6d81d372349 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

AdaVid: Adaptive Video-Language Pretraining Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.907252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.989080Z digest=sha256:059c45ef0769f9ce6ecf37153ae79d01c88a2a6187e0034b4144c43cd5574124

Observation 096fdd9b-bd77-4d49-a29c-37af99143617 · outbound

This paper cites Vision transformer slimming: Multi-dimension searching in continuous optimiza- tion space.

AdaVid: Adaptive Video-Language Pretraining Vision transformer slimming: Multi-dimension searching in continuous optimiza- tion space

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.892774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:55.993662Z digest=sha256:fb7f3fd0518de1a6053ac8fb4239effc7cb01c4409244fd919ab32f7388af8e4

Observation 0722223b-cbd0-402f-97d8-b8f9f094dccf · outbound

This paper cites Towards the Limit of Network Quantization.

AdaVid: Adaptive Video-Language Pretraining Towards the Limit of Network Quantization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:55.998203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:55.998203Z digest=sha256:e57ee63782ff5a4c5e0e540b38dae34932017c5ccaa42b52f4c4b8ef59f570c3

Observation 05a3fe94-49c0-4002-826b-fe1afda5c235 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

AdaVid: Adaptive Video-Language Pretraining BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.003372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.003372Z digest=sha256:65d4648f895e2e275f9944e93c2669a7629afd882cac50e85116b8c27e76cf31

Observation 440cb46e-30f8-4f78-8fa4-889d3d75a44a · outbound

This paper cites Matformer: Nested transformer for elastic inference.

AdaVid: Adaptive Video-Language Pretraining Matformer: Nested transformer for elastic inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.877891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.008261Z digest=sha256:604a231bac2844efafaaa04b422097fbaaaeea97345d0cbc66b59232abcb1fb1

Observation 77641918-6f5d-471a-85c3-986fcb0da2e0 · outbound

This paper cites Slowfast networks for video recognition.

AdaVid: Adaptive Video-Language Pretraining Slowfast networks for video recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.861335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.012750Z digest=sha256:ca176901799760325b7b055666a66b7bcafcfd701d6b8110112bd5ec608adc30

Observation 05e8a850-a9b9-4757-bc7a-853394a9cc16 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

AdaVid: Adaptive Video-Language Pretraining Ego4d: Around the world in 3,000 hours of egocentric video

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.846268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.017739Z digest=sha256:01d782883ab116613a68c600e60d3fd66d3647b9410b6d4dc597a1d12f06216e

Observation e3809cd4-506c-4386-91d2-8d6271a610d4 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

AdaVid: Adaptive Video-Language Pretraining Ego4d: Around the world in 3,000 hours of egocentric video

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.832277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.022435Z digest=sha256:e351e346eb7213dc6eef453b0c5e2b571e7c041996b727b4956cbc4b97c9d627

Observation da3a14b4-915a-41bd-924d-3e905235e210 · outbound

This paper cites Dynamic convnets on tiny devices via nested sparsity.

AdaVid: Adaptive Video-Language Pretraining Dynamic convnets on tiny devices via nested sparsity

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.818036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.026815Z digest=sha256:6bd51ac4ebf984c8dcd509adb1cd624987d914c0560ddfb4ab19697bb01f9bb1

Observation bcaa72c8-997d-445c-8839-ea0a049f07a6 · outbound

This paper cites Transkimmer: Transformer Learns to Layer-wise Skim.

AdaVid: Adaptive Video-Language Pretraining Transkimmer: Transformer Learns to Layer-wise Skim

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.031204Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.031204Z digest=sha256:dcc40f6feb13f5741ab478379ee6b725641448276df068d49a2a31cdd8351b2a

Observation fccdba26-2254-439b-892c-2250a6d5162c · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AdaVid: Adaptive Video-Language Pretraining Distilling the Knowledge in a Neural Network

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.035899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.035899Z digest=sha256:68e3b7f0c54096c86e12734978e2942242a5118c2f96e7436fb6974e755c34db

Observation 4640f7b3-90cc-446e-8e9c-ce294c95de7e · outbound

This paper cites Training Compute-Optimal Large Language Models.

AdaVid: Adaptive Video-Language Pretraining Training Compute-Optimal Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.040813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.040813Z digest=sha256:4fcff83d5bd97eb4fa3409b01164e57d8eace917123ec35f3c66ed4f72fc3257

Observation 10756d83-cb59-4617-841a-a84c00ef0b31 · outbound

This paper cites Dynabert: Dynamic bert with adaptive width and depth.

AdaVid: Adaptive Video-Language Pretraining Dynabert: Dynamic bert with adaptive width and depth

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.804032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.045191Z digest=sha256:9822191ffb4a151a6f6da951d3601fe6e3af47b57dd7dca92cb45b1572e3130f

Observation 7a036d8c-987c-43d3-a4b4-52b26a4e5795 · outbound

This paper cites Long movie clip classification with state-space video models.

AdaVid: Adaptive Video-Language Pretraining Long movie clip classification with state-space video models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.049562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.049562Z digest=sha256:8d1f3b2a3ab01230c0cf86a5b9e038d32e8c0682dc98d86ad93d89b21d3faaf6

Observation 3fbb887f-d92d-485c-ab64-82a673efe4b0 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

AdaVid: Adaptive Video-Language Pretraining Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.054262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.054262Z digest=sha256:d5f5cb612384bcdf80d8a17bb967f93b94cb2f01573142a6f682827a45d40aaf

Observation d14186c5-df10-40b1-a9b3-05541d0585c1 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

AdaVid: Adaptive Video-Language Pretraining Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.058930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.058930Z digest=sha256:8d5424986529f5cf34d9e6c64abb36da17608b5b5685a90fa41e4ba40d965fdb

Observation 0615abc7-7681-4708-a7c2-e681c661cba4 · outbound

This paper cites Matryoshka representation learning.

AdaVid: Adaptive Video-Language Pretraining Matryoshka representation learning

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.780186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.063785Z digest=sha256:5e4fb16732c098d15119ee8ccc1baf4de39a2304352577360a7d85551df817a8

Observation c354b6f8-641f-409a-b30b-0c17bbe8b3ee · outbound

This paper cites Block Pruning For Faster Transformers.

AdaVid: Adaptive Video-Language Pretraining Block Pruning For Faster Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.068793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.068793Z digest=sha256:b81506199e5d5bd0ab37b98492f11caeffa858a8a6d93832db1bd87c537e75e1

Observation cc3b5e32-5c90-4b00-be4f-a08a413218d0 · outbound

This paper cites Resound: Towards action recognition without representation bias.

AdaVid: Adaptive Video-Language Pretraining Resound: Towards action recognition without representation bias

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.765223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.073483Z digest=sha256:0424613b88b278719ecc45cbe4d3b98c6567084b1eadf64a5f47c48187547659

Observation ae65ca19-72ad-4e95-bdf9-61c8ab151262 · outbound

This paper cites Egocentric Video-Language Pretraining.

AdaVid: Adaptive Video-Language Pretraining Egocentric Video-Language Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.078363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.078363Z digest=sha256:4a9d0afdc70fdc9ff810a934161f98421c3559a02931b42d2d0e297e237bc9fb

Observation ce75f7db-64a8-4ce0-9e2e-526e496bf87d · outbound

This paper cites Egocentric video-language pretraining.

AdaVid: Adaptive Video-Language Pretraining Egocentric video-language pretraining

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.749579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.083071Z digest=sha256:dd5f203a66364291c2a8a95a68be0ec0845cabbb910d9117c5b35744225f3b5f

Observation c51ad084-7082-426e-b1d4-73f91270ff0b · outbound

This paper cites Rethinking the Value of Network Pruning.

AdaVid: Adaptive Video-Language Pretraining Rethinking the Value of Network Pruning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.087445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.087445Z digest=sha256:e714cddb91bce1bc43a73e6585dfe71995543f0687ba702d4ac34fd3bd27450b

Observation b0d0ef6f-8449-4594-a162-77d548a1b18a · outbound

This paper cites Decoupled Weight Decay Regularization.

AdaVid: Adaptive Video-Language Pretraining Decoupled Weight Decay Regularization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.092058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.092058Z digest=sha256:c20b1a72b78b006a998b287746523f747f68efd4f43243f37527fddc54f45c4f

Observation d2053245-e3a9-4b9c-93a3-30288afa8d6e · outbound

This paper cites UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation.

AdaVid: Adaptive Video-Language Pretraining UniVL: A Unified Video and Language Pre-Training Model for Multimodal Understanding and Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.096502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.096502Z digest=sha256:373a19eaeea5b7ec6818ceecef2163e39ca51485f6a3b538ef8163454ccbc60e

Observation 87c13dce-107c-453e-8ca7-06937970a679 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long- form video language understanding.

AdaVid: Adaptive Video-Language Pretraining Egoschema: A diagnostic benchmark for very long- form video language understanding

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.735087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.101319Z digest=sha256:8138d330b6a6fa1c7f725f2656346add7a1779bc977866336247d3e0c27150bc

Observation d2b4998c-b90c-492c-baa6-b8160e4604e6 · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips.

AdaVid: Adaptive Video-Language Pretraining Howto100m: Learning a text-video embedding by watching hundred mil- lion narrated video clips

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.720661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.105835Z digest=sha256:041cb43bce290178d7a47781dcb7c8ccc3bd0241c8819881325dbce742289433

Observation 8a0b218f-5f8f-4904-8cf0-e37b93acfa7b · outbound

This paper cites End-to-end learn- ing of visual representations from uncurated instructional videos.

AdaVid: Adaptive Video-Language Pretraining End-to-end learn- ing of visual representations from uncurated instructional videos

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.706236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.110318Z digest=sha256:b0831f3755467461ad8025c62296495fc16866c4dcca0472ae497142c9fa45f5

Observation cc4a0554-9109-4996-9e89-d03841bb85ca · outbound

This paper cites A sim- ple recipe for contrastively pre-training video-first encoders beyond 16 frames.

AdaVid: Adaptive Video-Language Pretraining A sim- ple recipe for contrastively pre-training video-first encoders beyond 16 frames

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.691703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.114688Z digest=sha256:19ae10938ceaa67b696f72b0111e01f16a4fb66a60e1c5e334b528cce80e6246

Observation 721ea63a-ec2c-4c0c-93db-3822f6387b3a · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AdaVid: Adaptive Video-Language Pretraining Learning transferable visual models from natural language supervi- sion

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.119502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.119502Z digest=sha256:40b628e67429cf0d94eb933f9b5f4cd6f211725293131032a0a165e51b79b13e

Observation bc351f54-06a2-4124-a937-65c6e3239472 · outbound

This paper cites SHARCS: Efficient Transformers through Routing with Dynamic Width Sub-networks.

AdaVid: Adaptive Video-Language Pretraining SHARCS: Efficient Transformers through Routing with Dynamic Width Sub-networks

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-16T12:35:56.322949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.124027Z digest=sha256:729b77d03fb176efb0f1e0ebc6668031b77c7e9987e74e842e34251a840f1f30

Observation a1a911d8-a3d9-4c5e-938e-8a01a045b54e · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

AdaVid: Adaptive Video-Language Pretraining DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.128650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.128650Z digest=sha256:3a732257ab650c11d5cf3122ee6bee32bf17e73e3b268cabd66648d6387e2ddf

Observation e87bcd01-65a0-40f3-874b-df2ac13f6d83 · outbound

This paper cites Q- bert: Hessian based ultra low precision quantization of bert.

AdaVid: Adaptive Video-Language Pretraining Q- bert: Hessian based ultra low precision quantization of bert

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.668159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.133416Z digest=sha256:07324050a8f7ae7f36cc4fa100e209e0d4cbd41e5d7ac296d22488ac851f5810

Observation 63612549-91f1-42c6-84b6-e2ea945978e7 · outbound

This paper cites Training data-efficient image transformers & distillation through atten- tion.

AdaVid: Adaptive Video-Language Pretraining Training data-efficient image transformers & distillation through atten- tion

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.653140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.138106Z digest=sha256:898f00e7738a4766339997c51a0471bcf1980a1857c6e861e073cc1fd54bb49f

Observation 2f04f4e1-1728-4516-9703-d2657e51f998 · outbound

This paper cites Attention is all you need.

AdaVid: Adaptive Video-Language Pretraining Attention is all you need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.142606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.142606Z digest=sha256:7b1027cdca7d730c770d647fe65ee43a9cda6390e4ef5ecf1476e9cff169dcfe

Observation d3a657a0-3aed-46bf-99ea-320548a525cc · outbound

This paper cites Deformable video trans- former.

AdaVid: Adaptive Video-Language Pretraining Deformable video trans- former

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.628578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.147182Z digest=sha256:f61221f5ac407936b4244a573eecbd4ca6c5f7f3aa49187c20efff892f65a45a

Observation 24561286-af50-4f6d-a761-985ccb83a42c · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

AdaVid: Adaptive Video-Language Pretraining VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.151992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.151992Z digest=sha256:4988adbefbdf17ce469140dd2358a13198b84992c6caf2c97fc33b624500d97d

Observation 4c970475-6e73-41c5-a6db-0b8dcfacdd20 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

AdaVid: Adaptive Video-Language Pretraining InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.157243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.157243Z digest=sha256:779a3535ac5c9087ed85b9e160d50737e13cadc3b5bac7a3fffadefc1e60d5d5

Observation 8060ef7f-6824-4eda-b5e7-a5e9ba352308 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

AdaVid: Adaptive Video-Language Pretraining Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.162838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.162838Z digest=sha256:2ebca9a0ac142e5d4610f6e71d9ffa07ed71af65849e3670ecac18719787e141

Observation 052cd404-e35f-47d8-b9a3-ced47168b153 · outbound

This paper cites VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding.

AdaVid: Adaptive Video-Language Pretraining VideoCLIP: Contrastive Pre-training for Zero-shot Video-Text Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.167427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.167427Z digest=sha256:ec5d452ec3efef42822bb20d1a1838096f86dcd87ea4d1f0b946b7614f0f9a28

Observation c6b921f7-1817-4f60-b8c9-178b349f5c63 · outbound

This paper cites Slimmable Neural Networks.

AdaVid: Adaptive Video-Language Pretraining Slimmable Neural Networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.172291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.172291Z digest=sha256:bbf0f9a15b112644ada55c1fea4a148767a896a69163753f35bdb44b1162699b

Observation c19320ce-cf02-4109-856d-7fe29833806a · outbound

This paper cites Self-chained image-language model for video localization and question answering.

AdaVid: Adaptive Video-Language Pretraining Self-chained image-language model for video localization and question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.177310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.177310Z digest=sha256:cf0b778451fd231f66d175245c960f75573a353ab80e02aa2fe8013cd060fc0a

Observation 5739293f-3212-42ca-a1a7-c8fe5a48e90b · outbound

This paper cites Poa: Pre-training once for models of all sizes.

AdaVid: Adaptive Video-Language Pretraining Poa: Pre-training once for models of all sizes

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.595138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.181965Z digest=sha256:0ad24bd11cb0fe9bfb91fe2ca98609ded79fdcde8de5b67bc2ec788cec2418f5

Observation d9093c44-da05-4731-ab8f-e1582d8778d9 · outbound

This paper cites Mgsampler: An explainable sampling strategy for video ac- tion recognition.

AdaVid: Adaptive Video-Language Pretraining Mgsampler: An explainable sampling strategy for video ac- tion recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-16T12:35:56.186656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:35:56.186656Z digest=sha256:02fb875fb5475e8abd0fe5796f2e680d31566bf111f6dc9d61a269f16c8c458b

Observation 907ccd42-b7ab-41b0-88fb-8defaf1949d5 · outbound

This paper cites FLOPs computation A.1.

AdaVid: Adaptive Video-Language Pretraining FLOPs computation A.1

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T12:35:56.571000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T12:35:56.191733Z digest=sha256:5eb9eda0b92c343b20f4d4d82f6a3e0d75e4923a386f1fd405c80bd92729e24b

Pith citing papers

No inbound Pith citation observations are available.